Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yuan, Chaohao, Xiao, Chenghao, Rong, Yu, Cheng, Hong, Huang, Long-Kai |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
IBCircuit: Towards Holistic Circuit Discovery with Information Bottleneck
par: Bian, Tian, et autres
Publié: (2026)
par: Bian, Tian, et autres
Publié: (2026)
ParaFormer: A Generalized PageRank Graph Transformer for Graph Representation Learning
par: Yuan, Chaohao, et autres
Publié: (2025)
par: Yuan, Chaohao, et autres
Publié: (2025)
Annotation-guided Protein Design with Multi-Level Domain Alignment
par: Yuan, Chaohao, et autres
Publié: (2024)
par: Yuan, Chaohao, et autres
Publié: (2024)
Decoupling Weighing and Selecting for Integrating Multiple Graph Pre-training Tasks
par: Fan, Tianyu, et autres
Publié: (2024)
par: Fan, Tianyu, et autres
Publié: (2024)
A Survey of Graph Transformers: Architectures, Theories and Applications
par: Yuan, Chaohao, et autres
Publié: (2025)
par: Yuan, Chaohao, et autres
Publié: (2025)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
par: Huang, Fanding, et autres
Publié: (2025)
par: Huang, Fanding, et autres
Publié: (2025)
Data-Efficient RLVR via Off-Policy Influence Guidance
par: Zhu, Erle, et autres
Publié: (2025)
par: Zhu, Erle, et autres
Publié: (2025)
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
par: Shu, Yao, et autres
Publié: (2026)
par: Shu, Yao, et autres
Publié: (2026)
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
par: Feng, Pengchao, et autres
Publié: (2025)
par: Feng, Pengchao, et autres
Publié: (2025)
Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
par: Xie, Pei-Xi, et autres
Publié: (2026)
par: Xie, Pei-Xi, et autres
Publié: (2026)
RLVR-World: Training World Models with Reinforcement Learning
par: Wu, Jialong, et autres
Publié: (2025)
par: Wu, Jialong, et autres
Publié: (2025)
Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR
par: Min, Zijun, et autres
Publié: (2026)
par: Min, Zijun, et autres
Publié: (2026)
ASD Classification on Dynamic Brain Connectome using Temporal Random Walk with Transformer-based Dynamic Network Embedding
par: Piriyasatit, Suchanuch, et autres
Publié: (2025)
par: Piriyasatit, Suchanuch, et autres
Publié: (2025)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
par: Hao, Zhezheng, et autres
Publié: (2025)
par: Hao, Zhezheng, et autres
Publié: (2025)
Learning Word Embedding with Better Distance Weighting and Window Size Scheduling
par: Yang, Chaohao, et autres
Publié: (2024)
par: Yang, Chaohao, et autres
Publié: (2024)
Pass@k Metric for RLVR: A Diagnostic Tool of Exploration, But Not an Objective
par: Yu, Yang
Publié: (2025)
par: Yu, Yang
Publié: (2025)
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
par: Koh, Woosung, et autres
Publié: (2026)
par: Koh, Woosung, et autres
Publié: (2026)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
par: Hu, Yuelin, et autres
Publié: (2026)
par: Hu, Yuelin, et autres
Publié: (2026)
Rapid Learning in Constrained Minimax Games with Negative Momentum
par: Fang, Zijian, et autres
Publié: (2024)
par: Fang, Zijian, et autres
Publié: (2024)
Unifying Structural Proximity and Equivalence for Enhanced Dynamic Network Embedding
par: Piriyasatit, Suchanuch, et autres
Publié: (2025)
par: Piriyasatit, Suchanuch, et autres
Publié: (2025)
Self-Distilled RLVR
par: Yang, Chenxu, et autres
Publié: (2026)
par: Yang, Chenxu, et autres
Publié: (2026)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
par: Zhao, Anhao, et autres
Publié: (2026)
par: Zhao, Anhao, et autres
Publié: (2026)
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
par: Shin, Haebin, et autres
Publié: (2025)
par: Shin, Haebin, et autres
Publié: (2025)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
par: Wang, Bo, et autres
Publié: (2025)
par: Wang, Bo, et autres
Publié: (2025)
The Path Not Taken: RLVR Provably Learns Off the Principals
par: Zhu, Hanqing, et autres
Publié: (2025)
par: Zhu, Hanqing, et autres
Publié: (2025)
Task Vector Quantization for Memory-Efficient Model Merging
par: Kim, Youngeun, et autres
Publié: (2025)
par: Kim, Youngeun, et autres
Publié: (2025)
Evaluating Parameter Efficient Methods for RLVR
par: Yin, Qingyu, et autres
Publié: (2025)
par: Yin, Qingyu, et autres
Publié: (2025)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
par: Chen, Zhipeng, et autres
Publié: (2026)
par: Chen, Zhipeng, et autres
Publié: (2026)
Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration
par: Zhao, Yang, et autres
Publié: (2026)
par: Zhao, Yang, et autres
Publié: (2026)
Whoever Started the Interference Should End It: Guiding Data-Free Model Merging via Task Vectors
par: Cheng, Runxi, et autres
Publié: (2025)
par: Cheng, Runxi, et autres
Publié: (2025)
Debunk the Myth of SFT Generalization
par: Lin, Xiaofeng, et autres
Publié: (2025)
par: Lin, Xiaofeng, et autres
Publié: (2025)
RLPR: Extrapolating RLVR to General Domains without Verifiers
par: Yu, Tianyu, et autres
Publié: (2025)
par: Yu, Tianyu, et autres
Publié: (2025)
Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
par: Fan, Chongyu, et autres
Publié: (2026)
par: Fan, Chongyu, et autres
Publié: (2026)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
par: Cui, Sijia, et autres
Publié: (2026)
par: Cui, Sijia, et autres
Publié: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
par: Wu, Junkang, et autres
Publié: (2025)
par: Wu, Junkang, et autres
Publié: (2025)
Not only where, But when: Temporal Scheduling for RLVR
par: Zhang, Jinghao, et autres
Publié: (2026)
par: Zhang, Jinghao, et autres
Publié: (2026)
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
par: Li, Kaiyuan, et autres
Publié: (2026)
par: Li, Kaiyuan, et autres
Publié: (2026)
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
par: Zhang, Xuechen, et autres
Publié: (2025)
par: Zhang, Xuechen, et autres
Publié: (2025)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
par: Ye, Hao, et autres
Publié: (2026)
par: Ye, Hao, et autres
Publié: (2026)
Rewards as Labels: Revisiting RLVR from a Classification Perspective
par: Zhai, Zepeng, et autres
Publié: (2026)
par: Zhai, Zepeng, et autres
Publié: (2026)
Documents similaires
-
IBCircuit: Towards Holistic Circuit Discovery with Information Bottleneck
par: Bian, Tian, et autres
Publié: (2026) -
ParaFormer: A Generalized PageRank Graph Transformer for Graph Representation Learning
par: Yuan, Chaohao, et autres
Publié: (2025) -
Annotation-guided Protein Design with Multi-Level Domain Alignment
par: Yuan, Chaohao, et autres
Publié: (2024) -
Decoupling Weighing and Selecting for Integrating Multiple Graph Pre-training Tasks
par: Fan, Tianyu, et autres
Publié: (2024) -
A Survey of Graph Transformers: Architectures, Theories and Applications
par: Yuan, Chaohao, et autres
Publié: (2025)