How to Mitigate Overfitting in Weak-to-strong Generalization?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Junhao, Cheng, Qinyuan, Fei, Zhaoye, Zheng, Yining, Guo, Qipeng, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
von: Shi, Junhao, et al.
Veröffentlicht: (2025)
von: Shi, Junhao, et al.
Veröffentlicht: (2025)
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
Mousse: Rectifying the Geometry of Muon with Curvature-Aware Preconditioning
von: Zhang, Yechen, et al.
Veröffentlicht: (2026)
von: Zhang, Yechen, et al.
Veröffentlicht: (2026)
Towards Mitigating Architecture Overfitting on Distilled Datasets
von: Zhong, Xuyang, et al.
Veröffentlicht: (2023)
von: Zhong, Xuyang, et al.
Veröffentlicht: (2023)
How to Set the Learning Rate for Large-Scale Pre-training?
von: Zhou, Yunhua, et al.
Veröffentlicht: (2026)
von: Zhou, Yunhua, et al.
Veröffentlicht: (2026)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
Mitigating Overfitting in Graph Neural Networks via Feature and Hyperplane Perturbation
von: Choi, Yoonhyuk, et al.
Veröffentlicht: (2022)
von: Choi, Yoonhyuk, et al.
Veröffentlicht: (2022)
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
von: Fei, Zhaoye, et al.
Veröffentlicht: (2025)
von: Fei, Zhaoye, et al.
Veröffentlicht: (2025)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
Provable Robust Overfitting Mitigation in Wasserstein Distributionally Robust Optimization
von: Liu, Shuang, et al.
Veröffentlicht: (2025)
von: Liu, Shuang, et al.
Veröffentlicht: (2025)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
Detecting Generative Parroting through Overfitting Masked Autoencoders
von: Taghanaki, Saeid Asgari, et al.
Veröffentlicht: (2024)
von: Taghanaki, Saeid Asgari, et al.
Veröffentlicht: (2024)
Control of Overfitting with Physics
von: Kozyrev, Sergei V., et al.
Veröffentlicht: (2024)
von: Kozyrev, Sergei V., et al.
Veröffentlicht: (2024)
How Ensemble Learning Balances Accuracy and Overfitting: A Bias-Variance Perspective on Tabular Data
von: Mohammad, Zubair Ahmed
Veröffentlicht: (2025)
von: Mohammad, Zubair Ahmed
Veröffentlicht: (2025)
TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
von: Fu, Lucheng, et al.
Veröffentlicht: (2026)
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
Benign Overfitting in Adversarial Training for Vision Transformers
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
Residual Stream Analysis of Overfitting And Structural Disruptions
von: Liu, Quan, et al.
Veröffentlicht: (2026)
von: Liu, Quan, et al.
Veröffentlicht: (2026)
Friend or Foe? Harnessing Controllable Overfitting for Anomaly Detection
von: Qian, Long, et al.
Veröffentlicht: (2024)
von: Qian, Long, et al.
Veröffentlicht: (2024)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
von: Ye, Hao, et al.
Veröffentlicht: (2026)
von: Ye, Hao, et al.
Veröffentlicht: (2026)
Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step Sizes
von: Qiao, Dan, et al.
Veröffentlicht: (2024)
von: Qiao, Dan, et al.
Veröffentlicht: (2024)
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
von: Guo, Yipin, et al.
Veröffentlicht: (2024)
von: Guo, Yipin, et al.
Veröffentlicht: (2024)
A Classical View on Benign Overfitting: The Role of Sample Size
von: Park, Junhyung, et al.
Veröffentlicht: (2025)
von: Park, Junhyung, et al.
Veröffentlicht: (2025)
LoRA Dropout as a Sparsity Regularizer for Overfitting Control
von: Lin, Yang, et al.
Veröffentlicht: (2024)
von: Lin, Yang, et al.
Veröffentlicht: (2024)
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
von: Ye, Qinyuan, et al.
Veröffentlicht: (2025)
von: Ye, Qinyuan, et al.
Veröffentlicht: (2025)
Overcoming Overfitting in Reinforcement Learning via Gaussian Process Diffusion Policy
von: Horprasert, Amornyos, et al.
Veröffentlicht: (2025)
von: Horprasert, Amornyos, et al.
Veröffentlicht: (2025)
From Overfitting to Reliability: Introducing the Hierarchical Approximate Bayesian Neural Network
von: Amirkhanian, Hayk, et al.
Veröffentlicht: (2025)
von: Amirkhanian, Hayk, et al.
Veröffentlicht: (2025)
Benign Overfitting for Regression with Trained Two-Layer ReLU Networks
von: Park, Junhyung, et al.
Veröffentlicht: (2024)
von: Park, Junhyung, et al.
Veröffentlicht: (2024)
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2026)
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
von: Zhao, Mengnan, et al.
Veröffentlicht: (2026)
von: Zhao, Mengnan, et al.
Veröffentlicht: (2026)
Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training
von: Zhang, Mozhi, et al.
Veröffentlicht: (2025)
von: Zhang, Mozhi, et al.
Veröffentlicht: (2025)
Risk Phase Transitions in Spiked Regression: Alignment Driven Benign and Catastrophic Overfitting
von: Li, Jiping, et al.
Veröffentlicht: (2025)
von: Li, Jiping, et al.
Veröffentlicht: (2025)
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
von: Koh, Woosung, et al.
Veröffentlicht: (2026)
von: Koh, Woosung, et al.
Veröffentlicht: (2026)
SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation
von: Cheng, Shuang, et al.
Veröffentlicht: (2025)
von: Cheng, Shuang, et al.
Veröffentlicht: (2025)
Understanding Hardness of Vision-Language Compositionality from A Token-level Causal Lens
von: Chen, Ziliang, et al.
Veröffentlicht: (2025)
von: Chen, Ziliang, et al.
Veröffentlicht: (2025)
Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
von: Jin, Luozhijie, et al.
Veröffentlicht: (2025)
von: Jin, Luozhijie, et al.
Veröffentlicht: (2025)
Agent Alignment in Evolving Social Norms
von: Li, Shimin, et al.
Veröffentlicht: (2024)
von: Li, Shimin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
von: Shi, Junhao, et al.
Veröffentlicht: (2025) -
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024) -
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024) -
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025) -
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
von: Wang, Bo, et al.
Veröffentlicht: (2025)