Gespeichert in:
| Hauptverfasser: | Wang, Hao, Gu, Hao, Piao, Hongming, Gong, Kaixiong, Ye, Yuxiao, Yue, Xiangyu, Han, Sirui, Guo, Yike, Wu, Dapeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.02244 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TinyThinker: Distilling Reasoning through Coarse-to-Fine Knowledge Internalization with Self-Reflection
von: Piao, Shengmin, et al.
Veröffentlicht: (2024)
von: Piao, Shengmin, et al.
Veröffentlicht: (2024)
Basic Reading Distillation
von: Zhou, Zhi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhi, et al.
Veröffentlicht: (2025)
Supervised Fine-Tuning as Inverse Reinforcement Learning
von: Sun, Hao
Veröffentlicht: (2024)
von: Sun, Hao
Veröffentlicht: (2024)
Rotation-Preserving Supervised Fine-Tuning
von: Jin, Hangzhan, et al.
Veröffentlicht: (2026)
von: Jin, Hangzhan, et al.
Veröffentlicht: (2026)
Staying Healthy While You Are Pregnant
Veröffentlicht: (2025)
Veröffentlicht: (2025)
On-Policy Supervised Fine-Tuning for Efficient Reasoning
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
von: Guo, Xin, et al.
Veröffentlicht: (2026)
von: Guo, Xin, et al.
Veröffentlicht: (2026)
Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs
von: Lu, Yuxiao, et al.
Veröffentlicht: (2024)
von: Lu, Yuxiao, et al.
Veröffentlicht: (2024)
Understanding Overadaptation in Supervised Fine-Tuning: The Role of Ensemble Methods
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
von: Diao, Muxi, et al.
Veröffentlicht: (2026)
von: Diao, Muxi, et al.
Veröffentlicht: (2026)
Closed-Loop Supervised Fine-Tuning of Tokenized Traffic Models
von: Zhang, Zhejun, et al.
Veröffentlicht: (2024)
von: Zhang, Zhejun, et al.
Veröffentlicht: (2024)
Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning
von: Goel, Jyotin, et al.
Veröffentlicht: (2026)
von: Goel, Jyotin, et al.
Veröffentlicht: (2026)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning
von: Li, Zhaoyi, et al.
Veröffentlicht: (2026)
von: Li, Zhaoyi, et al.
Veröffentlicht: (2026)
Video-R1: Reinforcing Video Reasoning in MLLMs
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
Preserving Diversity in Supervised Fine-Tuning of Large Language Models
von: Li, Ziniu, et al.
Veröffentlicht: (2024)
von: Li, Ziniu, et al.
Veröffentlicht: (2024)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
Supervised Fine-Tuning Needs to Unlock the Potential of Token Priority
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
BIFRÖST: 3D-Aware Image compositing with Language Instructions
von: Li, Lingxiao, et al.
Veröffentlicht: (2024)
von: Li, Lingxiao, et al.
Veröffentlicht: (2024)
How to Stay Curious while Avoiding Noisy TVs using Aleatoric Uncertainty Estimation
von: Mavor-Parker, Augustine N., et al.
Veröffentlicht: (2021)
von: Mavor-Parker, Augustine N., et al.
Veröffentlicht: (2021)
Fine-Tuning Robot Policies While Maintaining User Privacy
von: Christie, Benjamin A., et al.
Veröffentlicht: (2025)
von: Christie, Benjamin A., et al.
Veröffentlicht: (2025)
Preserving Multilingual Quality While Tuning Query Encoder on English Only
von: Vasilyev, Oleg, et al.
Veröffentlicht: (2024)
von: Vasilyev, Oleg, et al.
Veröffentlicht: (2024)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
von: Pang, Jinlong, et al.
Veröffentlicht: (2025)
von: Pang, Jinlong, et al.
Veröffentlicht: (2025)
InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
von: Jiang, Liming, et al.
Veröffentlicht: (2025)
von: Jiang, Liming, et al.
Veröffentlicht: (2025)
Policy Gradient with Adaptive Entropy Annealing for Continual Fine-Tuning
von: Zhang, Yaqian, et al.
Veröffentlicht: (2026)
von: Zhang, Yaqian, et al.
Veröffentlicht: (2026)
A Layer-wise Analysis of Supervised Fine-Tuning
von: Zhao, Qinghua, et al.
Veröffentlicht: (2026)
von: Zhao, Qinghua, et al.
Veröffentlicht: (2026)
Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis
von: Huang, Hong, et al.
Veröffentlicht: (2025)
von: Huang, Hong, et al.
Veröffentlicht: (2025)
EGAD: Entropy-Guided Adaptive Distillation for Token-Level Knowledge Transfer
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
Natural Language Fine-Tuning
von: Liu, Jia, et al.
Veröffentlicht: (2024)
von: Liu, Jia, et al.
Veröffentlicht: (2024)
Self-Supervised On-Policy Distillation for Reasoning Language Models
von: Tan, Zhiquan, et al.
Veröffentlicht: (2026)
von: Tan, Zhiquan, et al.
Veröffentlicht: (2026)
Pioneering Reliable Assessment in Text-to-Image Knowledge Editing: Leveraging a Fine-Grained Dataset and an Innovative Criterion
von: Gu, Hengrui, et al.
Veröffentlicht: (2024)
von: Gu, Hengrui, et al.
Veröffentlicht: (2024)
Reasoning While Recommending: Entropy-Guided Latent Reasoning in Generative Re-ranking Models
von: Zhang, Changshuo
Veröffentlicht: (2026)
von: Zhang, Changshuo
Veröffentlicht: (2026)
M-GRPO: Stabilizing Self-Supervised Reinforcement Learning for Large Language Models with Momentum-Anchored Policy Optimization
von: Bai, Bizhe, et al.
Veröffentlicht: (2025)
von: Bai, Bizhe, et al.
Veröffentlicht: (2025)
Remote Training in Task-Oriented Communication: Supervised or Self-Supervised with Fine-Tuning?
von: Li, Hongru, et al.
Veröffentlicht: (2025)
von: Li, Hongru, et al.
Veröffentlicht: (2025)
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
QFFT, Question-Free Fine-Tuning for Adaptive Reasoning
von: Liu, Wanlong, et al.
Veröffentlicht: (2025)
von: Liu, Wanlong, et al.
Veröffentlicht: (2025)
RoaD: Rollouts as Demonstrations for Closed-Loop Supervised Fine-Tuning of Autonomous Driving Policies
von: Garcia-Cobo, Guillermo, et al.
Veröffentlicht: (2025)
von: Garcia-Cobo, Guillermo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TinyThinker: Distilling Reasoning through Coarse-to-Fine Knowledge Internalization with Self-Reflection
von: Piao, Shengmin, et al.
Veröffentlicht: (2024) -
Basic Reading Distillation
von: Zhou, Zhi, et al.
Veröffentlicht: (2025) -
Supervised Fine-Tuning as Inverse Reinforcement Learning
von: Sun, Hao
Veröffentlicht: (2024) -
Rotation-Preserving Supervised Fine-Tuning
von: Jin, Hangzhan, et al.
Veröffentlicht: (2026) -
Staying Healthy While You Are Pregnant
Veröffentlicht: (2025)