Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Miaosen, Liu, Yishan, Lin, Shuxia, Yang, Xu, Dai, Qi, Luo, Chong, Jiang, Weihao, Hou, Peng, Zeng, Anxiang, Geng, Xin, Guo, Baining |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MageBench: Bridging Large Multimodal Models to Agents
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024)
Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024)
CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning
von: Tang, Zhijiang, et al.
Veröffentlicht: (2026)
von: Tang, Zhijiang, et al.
Veröffentlicht: (2026)
Improved Noise Schedule for Diffusion Training
von: Hang, Tiankai, et al.
Veröffentlicht: (2024)
von: Hang, Tiankai, et al.
Veröffentlicht: (2024)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
Transformer-empowered Multi-modal Item Embedding for Enhanced Image Search in E-Commerce
von: Liu, Chang, et al.
Veröffentlicht: (2023)
von: Liu, Chang, et al.
Veröffentlicht: (2023)
Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment Analysis
von: Luo, Miaosen, et al.
Veröffentlicht: (2025)
von: Luo, Miaosen, et al.
Veröffentlicht: (2025)
Controlled LLM Training on Spectral Sphere
von: Xie, Tian, et al.
Veröffentlicht: (2026)
von: Xie, Tian, et al.
Veröffentlicht: (2026)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation
von: Nie, Dong
Veröffentlicht: (2026)
von: Nie, Dong
Veröffentlicht: (2026)
Covering Human Action Space for Computer Use: Data Synthesis and Benchmark
von: Zhang, Miaosen, et al.
Veröffentlicht: (2026)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2026)
InfoAgent: Advancing Autonomous Information-Seeking Agents
von: Zhang, Gongrui, et al.
Veröffentlicht: (2025)
von: Zhang, Gongrui, et al.
Veröffentlicht: (2025)
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment
von: Li, Jiaxiang, et al.
Veröffentlicht: (2024)
von: Li, Jiaxiang, et al.
Veröffentlicht: (2024)
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT
von: Wang, Jiacheng, et al.
Veröffentlicht: (2026)
von: Wang, Jiacheng, et al.
Veröffentlicht: (2026)
ESPO: Entropy Importance Sampling Policy Optimization
von: Sheng, Yuepeng, et al.
Veröffentlicht: (2025)
von: Sheng, Yuepeng, et al.
Veröffentlicht: (2025)
Crowd-SFT: Crowdsourcing for LLM Alignment
von: Sotiropoulos, Alex, et al.
Veröffentlicht: (2025)
von: Sotiropoulos, Alex, et al.
Veröffentlicht: (2025)
From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
von: Li, Youpeng, et al.
Veröffentlicht: (2026)
von: Li, Youpeng, et al.
Veröffentlicht: (2026)
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
von: Qian, Wenxu, et al.
Veröffentlicht: (2025)
von: Qian, Wenxu, et al.
Veröffentlicht: (2025)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
Efficient Diffusion Training via Min-SNR Weighting Strategy
von: Hang, Tiankai, et al.
Veröffentlicht: (2023)
von: Hang, Tiankai, et al.
Veröffentlicht: (2023)
Pre-Trained Policy Discriminators are General Reward Models
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
CCA: Collaborative Competitive Agents for Image Editing
von: Hang, Tiankai, et al.
Veröffentlicht: (2024)
von: Hang, Tiankai, et al.
Veröffentlicht: (2024)
MUG-V 10B: High-efficiency Training Pipeline for Large Video Generation Models
von: Zhang, Yongshun, et al.
Veröffentlicht: (2025)
von: Zhang, Yongshun, et al.
Veröffentlicht: (2025)
RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents
von: Zhu, Jialiang, et al.
Veröffentlicht: (2026)
von: Zhu, Jialiang, et al.
Veröffentlicht: (2026)
Addressing Skewed Heterogeneity via Federated Prototype Rectification with Personalization
von: Guo, Shunxin, et al.
Veröffentlicht: (2024)
von: Guo, Shunxin, et al.
Veröffentlicht: (2024)
STHFL: Spatio-Temporal Heterogeneous Federated Learning
von: Guo, Shunxin, et al.
Veröffentlicht: (2025)
von: Guo, Shunxin, et al.
Veröffentlicht: (2025)
Microbial Memory of Drought Reshapes Root‐Associated Communities to Enhance Plant Resilience
von: Hongyin Qi, et al.
Veröffentlicht: (2025)
von: Hongyin Qi, et al.
Veröffentlicht: (2025)
Weak solutions to the steady incompressible Euler equations with source terms
von: Huang, Anxiang
Veröffentlicht: (2024)
von: Huang, Anxiang
Veröffentlicht: (2024)
Weak solutions to the steady compressible Euler equations with source terms
von: Huang, Anxiang
Veröffentlicht: (2024)
von: Huang, Anxiang
Veröffentlicht: (2024)
A Probabilistic Framework for Temporal Distribution Generalization in Industry-Scale Recommender Systems
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Towards Reliable Evaluation of Large Language Models for Multilingual and Multimodal E-Commerce Applications
von: Xie, Shuyi, et al.
Veröffentlicht: (2025)
von: Xie, Shuyi, et al.
Veröffentlicht: (2025)
Does the Question Really Matter? Training-Free Data Selection for Vision-Language SFT
von: Sun, Peng, et al.
Veröffentlicht: (2026)
von: Sun, Peng, et al.
Veröffentlicht: (2026)
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning
von: Zhu, Taojie, et al.
Veröffentlicht: (2026)
von: Zhu, Taojie, et al.
Veröffentlicht: (2026)
TMS: Trajectory-Mixed Supervision for Reward-Free, On-Policy SFT
von: Khan, Rana Muhammad Shahroz, et al.
Veröffentlicht: (2026)
von: Khan, Rana Muhammad Shahroz, et al.
Veröffentlicht: (2026)
Contract manufacturer versus platform: quality innovation strategy with brand spillover effect
von: Ling Li, et al.
Veröffentlicht: (2026)
von: Ling Li, et al.
Veröffentlicht: (2026)
Enhanced Sparsification via Stimulative Training
von: Tang, Shengji, et al.
Veröffentlicht: (2024)
von: Tang, Shengji, et al.
Veröffentlicht: (2024)
Language-Guided Face Animation by Recurrent StyleGAN-based Generator
von: Hang, Tiankai, et al.
Veröffentlicht: (2022)
von: Hang, Tiankai, et al.
Veröffentlicht: (2022)
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
von: Chen, Jierun, et al.
Veröffentlicht: (2025)
von: Chen, Jierun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MageBench: Bridging Large Multimodal Models to Agents
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024) -
Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024) -
CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning
von: Tang, Zhijiang, et al.
Veröffentlicht: (2026) -
Improved Noise Schedule for Diffusion Training
von: Hang, Tiankai, et al.
Veröffentlicht: (2024) -
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)