Understanding Overadaptation in Supervised Fine-Tuning: The Role of Ensemble Methods
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hao, Yifan, Pan, Xingyuan, Zhang, Hanning, Ye, Chenlu, Pan, Rui, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
Self-rewarding correction for mathematical reasoning
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
PhysProver: Advancing Automatic Theorem Proving for Physics
von: Zhang, Hanning, et al.
Veröffentlicht: (2026)
von: Zhang, Hanning, et al.
Veröffentlicht: (2026)
Towards Better Generalization via Distributional Input Projection Network
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
On-Policy Supervised Fine-Tuning for Efficient Reasoning
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
LR-SQL: A Supervised Fine-Tuning Method for Text2SQL Tasks under Low-Resource Scenarios
von: Wuzhenghong, Wen, et al.
Veröffentlicht: (2024)
von: Wuzhenghong, Wen, et al.
Veröffentlicht: (2024)
QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals
von: Zhang, Nan, et al.
Veröffentlicht: (2026)
von: Zhang, Nan, et al.
Veröffentlicht: (2026)
Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models
von: Jiang, Haitao, et al.
Veröffentlicht: (2026)
von: Jiang, Haitao, et al.
Veröffentlicht: (2026)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
von: Pan, Rui, et al.
Veröffentlicht: (2024)
von: Pan, Rui, et al.
Veröffentlicht: (2024)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
von: Ye, Chenlu, et al.
Veröffentlicht: (2026)
von: Ye, Chenlu, et al.
Veröffentlicht: (2026)
Proximal Supervised Fine-Tuning
von: Zhu, Wenhong, et al.
Veröffentlicht: (2025)
von: Zhu, Wenhong, et al.
Veröffentlicht: (2025)
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
Supervised Fine-Tuning as Inverse Reinforcement Learning
von: Sun, Hao
Veröffentlicht: (2024)
von: Sun, Hao
Veröffentlicht: (2024)
Adaptive Ensembles of Fine-Tuned Transformers for LLM-Generated Text Detection
von: Lai, Zhixin, et al.
Veröffentlicht: (2024)
von: Lai, Zhixin, et al.
Veröffentlicht: (2024)
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
von: Xia, Yuchen, et al.
Veröffentlicht: (2024)
von: Xia, Yuchen, et al.
Veröffentlicht: (2024)
TAGCOS: Task-agnostic Gradient Clustered Coreset Selection for Instruction Tuning Data
von: Zhang, Jipeng, et al.
Veröffentlicht: (2024)
von: Zhang, Jipeng, et al.
Veröffentlicht: (2024)
A Mechanistic Investigation of Supervised Fine Tuning
von: Chopra, Ruhaan
Veröffentlicht: (2026)
von: Chopra, Ruhaan
Veröffentlicht: (2026)
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
von: Pang, Jinlong, et al.
Veröffentlicht: (2025)
von: Pang, Jinlong, et al.
Veröffentlicht: (2025)
Supervised Fine-Tuning Needs to Unlock the Potential of Token Priority
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
Skillful High-Resolution Ensemble Precipitation Forecasting with an Integrated Deep Learning Framework
von: He, Shuangshuang, et al.
Veröffentlicht: (2025)
von: He, Shuangshuang, et al.
Veröffentlicht: (2025)
Rotation-Preserving Supervised Fine-Tuning
von: Jin, Hangzhan, et al.
Veröffentlicht: (2026)
von: Jin, Hangzhan, et al.
Veröffentlicht: (2026)
GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
von: Wang, Chun, et al.
Veröffentlicht: (2025)
von: Wang, Chun, et al.
Veröffentlicht: (2025)
Understanding and Preserving Safety in Fine-Tuned LLMs
von: Zhang, Jiawen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiawen, et al.
Veröffentlicht: (2026)
Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
Semi-Supervised Diseased Detection from Speech Dialogues with Multi-Level Data Modeling
von: Li, Xingyuan, et al.
Veröffentlicht: (2026)
von: Li, Xingyuan, et al.
Veröffentlicht: (2026)
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
Bayesian Natural Gradient Fine-Tuning of CLIP Models via Kalman Filtering
von: Abdi, Hossein, et al.
Veröffentlicht: (2025)
von: Abdi, Hossein, et al.
Veröffentlicht: (2025)
Efficient and Scalable Fine-Tune of Language Models for Genome Understanding
von: Zhan, Huixin, et al.
Veröffentlicht: (2024)
von: Zhan, Huixin, et al.
Veröffentlicht: (2024)
Addressing Bias Through Ensemble Learning and Regularized Fine-Tuning
von: Radwan, Ahmed, et al.
Veröffentlicht: (2024)
von: Radwan, Ahmed, et al.
Veröffentlicht: (2024)
A Graph Prompt Fine-Tuning Method for WSN Spatio-Temporal Correlation Anomaly Detection
von: Ye, Miao, et al.
Veröffentlicht: (2026)
von: Ye, Miao, et al.
Veröffentlicht: (2026)
GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification
von: Gan, Wangjie, et al.
Veröffentlicht: (2026)
von: Gan, Wangjie, et al.
Veröffentlicht: (2026)
FedPFT: Federated Proxy Fine-Tuning of Foundation Models
von: Peng, Zhaopeng, et al.
Veröffentlicht: (2024)
von: Peng, Zhaopeng, et al.
Veröffentlicht: (2024)
CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models
von: Han, Kairong, et al.
Veröffentlicht: (2025)
von: Han, Kairong, et al.
Veröffentlicht: (2025)
Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models
von: Li, Yuan, et al.
Veröffentlicht: (2025)
von: Li, Yuan, et al.
Veröffentlicht: (2025)
A Layer-wise Analysis of Supervised Fine-Tuning
von: Zhao, Qinghua, et al.
Veröffentlicht: (2026)
von: Zhao, Qinghua, et al.
Veröffentlicht: (2026)
Goal-Conditioned Supervised Learning for LLM Fine-Tuning
von: Li, Shijun, et al.
Veröffentlicht: (2026)
von: Li, Shijun, et al.
Veröffentlicht: (2026)
CharTool: Tool-Integrated Visual Reasoning for Chart Understanding
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
Zeroth-Order Fine-Tuning of LLMs in Random Subspaces
von: Yu, Ziming, et al.
Veröffentlicht: (2024)
von: Yu, Ziming, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models
von: Hao, Yifan, et al.
Veröffentlicht: (2025) -
Self-rewarding correction for mathematical reasoning
von: Xiong, Wei, et al.
Veröffentlicht: (2025) -
PhysProver: Advancing Automatic Theorem Proving for Physics
von: Zhang, Hanning, et al.
Veröffentlicht: (2026) -
Towards Better Generalization via Distributional Input Projection Network
von: Hao, Yifan, et al.
Veröffentlicht: (2025) -
On-Policy Supervised Fine-Tuning for Efficient Reasoning
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)