Saved in:
| Main Authors: | Gu, Weizheng, Li, Chengze, Yu, Zhuohao, Sun, Mengyuan, Yang, Zhibang, Wang, Wei, Jia, Hongrui, Zhang, Shikun, Ye, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.01611 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SteerRM: Debiasing Reward Models via Sparse Autoencoders
by: Sun, Mengyuan, et al.
Published: (2026)
by: Sun, Mengyuan, et al.
Published: (2026)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
by: Yu, Zhuohao, et al.
Published: (2025)
by: Yu, Zhuohao, et al.
Published: (2025)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
RewardAnything: Generalizable Principle-Following Reward Models
by: Yu, Zhuohao, et al.
Published: (2025)
by: Yu, Zhuohao, et al.
Published: (2025)
Learning Structure-Semantic Evolution Trajectories for Graph Domain Adaptation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
TMS: Trajectory-Mixed Supervision for Reward-Free, On-Policy SFT
by: Khan, Rana Muhammad Shahroz, et al.
Published: (2026)
by: Khan, Rana Muhammad Shahroz, et al.
Published: (2026)
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
by: Wu, Yongliang, et al.
Published: (2025)
by: Wu, Yongliang, et al.
Published: (2025)
Hyper-STTN: Hypergraph Augmented Spatial-Temporal Transformer Network for Trajectory Prediction
by: Wang, Weizheng, et al.
Published: (2024)
by: Wang, Weizheng, et al.
Published: (2024)
PepThink-R1: LLM for Interpretable Cyclic Peptide Optimization with CoT SFT and Reinforcement Learning
by: Wang, Ruheng, et al.
Published: (2025)
by: Wang, Ruheng, et al.
Published: (2025)
Mitigating Spurious Correlations with Causal Logit Perturbation
by: Zhou, Xiaoling, et al.
Published: (2025)
by: Zhou, Xiaoling, et al.
Published: (2025)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
What Do Latent Action Models Actually Learn?
by: Zhang, Chuheng, et al.
Published: (2025)
by: Zhang, Chuheng, et al.
Published: (2025)
ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models
by: Li, Chengze, et al.
Published: (2026)
by: Li, Chengze, et al.
Published: (2026)
Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
by: Kong, Deyang, et al.
Published: (2025)
by: Kong, Deyang, et al.
Published: (2025)
SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models
by: Kim, Gyuhak, et al.
Published: (2025)
by: Kim, Gyuhak, et al.
Published: (2025)
Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data
by: Jia, Zeyu, et al.
Published: (2024)
by: Jia, Zeyu, et al.
Published: (2024)
An update to PYRO-NN: A Python Library for Differentiable CT Operators
by: Schneider, Linda-Sophie, et al.
Published: (2025)
by: Schneider, Linda-Sophie, et al.
Published: (2025)
Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework
by: Jia, Hongrui, et al.
Published: (2025)
by: Jia, Hongrui, et al.
Published: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT
by: Wang, Jiacheng, et al.
Published: (2026)
by: Wang, Jiacheng, et al.
Published: (2026)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
by: Yu, Zhuohao, et al.
Published: (2026)
by: Yu, Zhuohao, et al.
Published: (2026)
Debunk the Myth of SFT Generalization
by: Lin, Xiaofeng, et al.
Published: (2025)
by: Lin, Xiaofeng, et al.
Published: (2025)
Bridging Global Intent with Local Details: A Hierarchical Representation Approach for Semantic Validation in Text-to-SQL
by: Qiu, Rihong, et al.
Published: (2025)
by: Qiu, Rihong, et al.
Published: (2025)
Procedural-skill SFT across capacity tiers: A W-Shaped pre-SFT Trajectory and Regime-Asymmetric Mechanism on 0.8B-4B Qwen3.5 Models
by: Strozzi, Igor
Published: (2026)
by: Strozzi, Igor
Published: (2026)
Spectral Heterogeneous Graph Convolutions via Positive Noncommutative Polynomials
by: He, Mingguo, et al.
Published: (2023)
by: He, Mingguo, et al.
Published: (2023)
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
by: Koh, Woosung, et al.
Published: (2026)
by: Koh, Woosung, et al.
Published: (2026)
Learning Wavelet-Sparse FDK for 3D Cone-Beam CT Reconstruction
by: Sun, Yipeng, et al.
Published: (2025)
by: Sun, Yipeng, et al.
Published: (2025)
Task-Oriented Low-Label Semantic Communication With Self-Supervised Learning
by: Gu, Run, et al.
Published: (2025)
by: Gu, Run, et al.
Published: (2025)
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
by: Lu, Aojun, et al.
Published: (2026)
by: Lu, Aojun, et al.
Published: (2026)
PatchAD: A Lightweight Patch-based MLP-Mixer for Time Series Anomaly Detection
by: Zhong, Zhijie, et al.
Published: (2024)
by: Zhong, Zhijie, et al.
Published: (2024)
Model Generalization on Text Attribute Graphs: Principles with Large Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Continual SFT Matches Multimodal RLHF with Negative Supervision
by: Zhu, Ke, et al.
Published: (2024)
by: Zhu, Ke, et al.
Published: (2024)
Learning by Doing: An Online Causal Reinforcement Learning Framework with Causal-Aware Policy
by: Cai, Ruichu, et al.
Published: (2024)
by: Cai, Ruichu, et al.
Published: (2024)
Enhancing In-Context Learning via Implicit Demonstration Augmentation
by: Zhou, Xiaoling, et al.
Published: (2024)
by: Zhou, Xiaoling, et al.
Published: (2024)
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning
by: Zhu, Taojie, et al.
Published: (2026)
by: Zhu, Taojie, et al.
Published: (2026)
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
by: Liu, Zihan, et al.
Published: (2025)
by: Liu, Zihan, et al.
Published: (2025)
PLATONT: Learning a Platonic Representation for Unified Network Tomography
by: Du, Chengze, et al.
Published: (2025)
by: Du, Chengze, et al.
Published: (2025)
Logarithmic Regret for Online KL-Regularized Reinforcement Learning
by: Zhao, Heyang, et al.
Published: (2025)
by: Zhao, Heyang, et al.
Published: (2025)
Similar Items
-
SteerRM: Debiasing Reward Models via Sparse Autoencoders
by: Sun, Mengyuan, et al.
Published: (2026) -
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
by: Yu, Zhuohao, et al.
Published: (2025) -
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
by: Yu, Zhuohao, et al.
Published: (2024) -
RewardAnything: Generalizable Principle-Following Reward Models
by: Yu, Zhuohao, et al.
Published: (2025) -
Learning Structure-Semantic Evolution Trajectories for Graph Domain Adaptation
by: Chen, Wei, et al.
Published: (2026)