Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Yongcan, He, Lingxiao, Lu, Shuo, Sheng, Lijun, Xu, Yinuo, Wang, Yanbo, Guo, Kuangpu, Cheng, Jianjie, Wang, Meng, Xie, Qianlong, Wang, Xingxing, Hu, Dapeng, Liang, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
by: Yu, Yongcan, et al.
Published: (2026)
by: Yu, Yongcan, et al.
Published: (2026)
How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1
by: Xu, Yinuo, et al.
Published: (2026)
by: Xu, Yinuo, et al.
Published: (2026)
DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing
by: Lu, Shuo, et al.
Published: (2025)
by: Lu, Shuo, et al.
Published: (2025)
Cooperative Pseudo Labeling for Unsupervised Federated Classification
by: Guo, Kuangpu, et al.
Published: (2025)
by: Guo, Kuangpu, et al.
Published: (2025)
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
by: Lu, Shuo, et al.
Published: (2026)
by: Lu, Shuo, et al.
Published: (2026)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
by: Yu, Yongcan, et al.
Published: (2025)
by: Yu, Yongcan, et al.
Published: (2025)
Proximal Supervised Fine-Tuning
by: Zhu, Wenhong, et al.
Published: (2025)
by: Zhu, Wenhong, et al.
Published: (2025)
On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning
by: Li, Zhaoyi, et al.
Published: (2026)
by: Li, Zhaoyi, et al.
Published: (2026)
STAMP: Outlier-Aware Test-Time Adaptation with Stable Memory Replay
by: Yu, Yongcan, et al.
Published: (2024)
by: Yu, Yongcan, et al.
Published: (2024)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
by: Wang, Yanbo, et al.
Published: (2026)
by: Wang, Yanbo, et al.
Published: (2026)
Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models
by: Ruan, Zhiwen, et al.
Published: (2026)
by: Ruan, Zhiwen, et al.
Published: (2026)
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative Ranks
by: Yang, Zhichao, et al.
Published: (2026)
by: Yang, Zhichao, et al.
Published: (2026)
Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapter
by: Jiang, Bo, et al.
Published: (2025)
by: Jiang, Bo, et al.
Published: (2025)
GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
by: Fiaz, Mustansar, et al.
Published: (2025)
by: Fiaz, Mustansar, et al.
Published: (2025)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
by: Tang, Zhengyang, et al.
Published: (2024)
by: Tang, Zhengyang, et al.
Published: (2024)
Knowledge Graph-Infused Fine-Tuning for Structured Reasoning in Large Language Models
by: Zhang, Wuyang, et al.
Published: (2025)
by: Zhang, Wuyang, et al.
Published: (2025)
Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging
by: Guo, Kuangpu, et al.
Published: (2025)
by: Guo, Kuangpu, et al.
Published: (2025)
Personalized Federated Learning via Dual-Prompt Optimization and Cross Fusion
by: Zhang, Yuguang, et al.
Published: (2025)
by: Zhang, Yuguang, et al.
Published: (2025)
FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training
by: Yu, Hongzhou, et al.
Published: (2025)
by: Yu, Hongzhou, et al.
Published: (2025)
SED-SFT: Selectively Encouraging Diversity in Supervised Fine-Tuning
by: Chen, Yijie, et al.
Published: (2026)
by: Chen, Yijie, et al.
Published: (2026)
Is Fine-Tuning an Effective Solution? Reassessing Knowledge Editing for Unstructured Data
by: Xiong, Hao, et al.
Published: (2025)
by: Xiong, Hao, et al.
Published: (2025)
R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
by: Sheng, Lijun, et al.
Published: (2025)
by: Sheng, Lijun, et al.
Published: (2025)
Towards Compatible Fine-tuning for Vision-Language Model Updates
by: Wang, Zhengbo, et al.
Published: (2024)
by: Wang, Zhengbo, et al.
Published: (2024)
Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization
by: Liang, Jian, et al.
Published: (2023)
by: Liang, Jian, et al.
Published: (2023)
Self-Tuning Self-Supervised Image Anomaly Detection
by: Yoo, Jaemin, et al.
Published: (2023)
by: Yoo, Jaemin, et al.
Published: (2023)
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
by: Fu, Yuqian, et al.
Published: (2025)
by: Fu, Yuqian, et al.
Published: (2025)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
by: He, Junwen, et al.
Published: (2024)
by: He, Junwen, et al.
Published: (2024)
DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
by: Ma, Yiming, et al.
Published: (2026)
by: Ma, Yiming, et al.
Published: (2026)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026)
by: Ma, Xueqi, et al.
Published: (2026)
Fine-Grained Motion Compression and Selective Temporal Fusion for Neural B-Frame Video Coding
by: Sheng, Xihua, et al.
Published: (2025)
by: Sheng, Xihua, et al.
Published: (2025)
QFFT, Question-Free Fine-Tuning for Adaptive Reasoning
by: Liu, Wanlong, et al.
Published: (2025)
by: Liu, Wanlong, et al.
Published: (2025)
Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models
by: Yu, Bin, et al.
Published: (2025)
by: Yu, Bin, et al.
Published: (2025)
Generative Large-Scale Pre-trained Models for Automated Ad Bidding Optimization
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
Stabilizing LLM Supervised Fine-Tuning via Explicit Distributional Control
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
by: Ye, Junjie, et al.
Published: (2025)
by: Ye, Junjie, et al.
Published: (2025)
A Survey on Data Selection for LLM Instruction Tuning
by: Zhang, Bolin, et al.
Published: (2024)
by: Zhang, Bolin, et al.
Published: (2024)
Similar Items
-
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
by: Yu, Yongcan, et al.
Published: (2026) -
How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1
by: Xu, Yinuo, et al.
Published: (2026) -
DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing
by: Lu, Shuo, et al.
Published: (2025) -
Cooperative Pseudo Labeling for Unsupervised Federated Classification
by: Guo, Kuangpu, et al.
Published: (2025) -
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
by: Wang, Yanbo, et al.
Published: (2025)