Hyperparameter Loss Surfaces Are Simple Near their Optima
Fuente:
arXiv
Saved in:
| Main Authors: | Lourie, Nicholas, He, He, Cho, Kyunghyun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Show Your Work with Confidence: Confidence Bands for Tuning Curves
by: Lourie, Nicholas, et al.
Published: (2023)
by: Lourie, Nicholas, et al.
Published: (2023)
Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
by: Lourie, Nicholas, et al.
Published: (2025)
by: Lourie, Nicholas, et al.
Published: (2025)
Neural Neural Scaling Laws
by: Hu, Michael Y., et al.
Published: (2026)
by: Hu, Michael Y., et al.
Published: (2026)
Aioli: A Unified Optimization Framework for Language Model Data Mixing
by: Chen, Mayee F., et al.
Published: (2024)
by: Chen, Mayee F., et al.
Published: (2024)
Hyperparameters in Continual Learning: A Reality Check
by: Cha, Sungmin, et al.
Published: (2024)
by: Cha, Sungmin, et al.
Published: (2024)
Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time
by: Hu, Michael Y., et al.
Published: (2026)
by: Hu, Michael Y., et al.
Published: (2026)
None To Optima in Few Shots: Bayesian Optimization with MDP Priors
by: Li, Diantong, et al.
Published: (2025)
by: Li, Diantong, et al.
Published: (2025)
On the Relationship Between the Choice of Representation and In-Context Learning
by: Marinescu, Ioana, et al.
Published: (2025)
by: Marinescu, Ioana, et al.
Published: (2025)
Efficient semantic uncertainty quantification in language models via diversity-steered sampling
by: Park, Ji Won, et al.
Published: (2025)
by: Park, Ji Won, et al.
Published: (2025)
Temporal Generalization: A Reality Check
by: Madaan, Divyam, et al.
Published: (2025)
by: Madaan, Divyam, et al.
Published: (2025)
Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable Modeling
by: Madaan, Divyam, et al.
Published: (2026)
by: Madaan, Divyam, et al.
Published: (2026)
Language Models as Causal Effect Generators
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
by: Madaan, Divyam, et al.
Published: (2025)
by: Madaan, Divyam, et al.
Published: (2025)
Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal Learning
by: Madaan, Divyam, et al.
Published: (2024)
by: Madaan, Divyam, et al.
Published: (2024)
On the Loss of Context-awareness in General Instruction Fine-tuning
by: Wang, Yihan, et al.
Published: (2024)
by: Wang, Yihan, et al.
Published: (2024)
SimpleGPT: Improving GPT via A Simple Normalization Strategy
by: Chen, Marco, et al.
Published: (2026)
by: Chen, Marco, et al.
Published: (2026)
SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters
by: Xiao, Teng, et al.
Published: (2025)
by: Xiao, Teng, et al.
Published: (2025)
Decoding Decoded: Understanding Hyperparameter Effects in Open-Ended Text Generation
by: Arias, Esteban Garces, et al.
Published: (2024)
by: Arias, Esteban Garces, et al.
Published: (2024)
Interim Report on Human-Guided Adaptive Hyperparameter Optimization with Multi-Fidelity Sprints
by: Kamfonas, Michael
Published: (2025)
by: Kamfonas, Michael
Published: (2025)
Translating Hanja Historical Documents to Contemporary Korean and English
by: Son, Juhee, et al.
Published: (2022)
by: Son, Juhee, et al.
Published: (2022)
Exploring Public Attention in the Circular Economy through Topic Modelling with Twin Hyperparameter Optimisation
by: Song, Junhao, et al.
Published: (2024)
by: Song, Junhao, et al.
Published: (2024)
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
by: Zeng, Weihao, et al.
Published: (2025)
by: Zeng, Weihao, et al.
Published: (2025)
Preference Learning Algorithms Do Not Learn Preference Rankings
by: Chen, Angelica, et al.
Published: (2024)
by: Chen, Angelica, et al.
Published: (2024)
Training Language Models with Language Feedback at Scale
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Optimizing Retrieval-Augmented Generation: Analysis of Hyperparameter Impact on Performance and Efficiency
by: Ammar, Adel, et al.
Published: (2025)
by: Ammar, Adel, et al.
Published: (2025)
Examining the Robustness of Homogeneity Bias to Hyperparameter Adjustments in GPT-4
by: Lee, Messi H. J.
Published: (2025)
by: Lee, Messi H. J.
Published: (2025)
The Unlearnability Phenomenon in RLVR for Language Models
by: Chen, Yulin, et al.
Published: (2026)
by: Chen, Yulin, et al.
Published: (2026)
ESURF: Simple and Effective EDU Segmentation
by: Sediqin, Mohammadreza, et al.
Published: (2025)
by: Sediqin, Mohammadreza, et al.
Published: (2025)
Simple and Effective Input Reformulations for Translation
by: Yu, Brian, et al.
Published: (2023)
by: Yu, Brian, et al.
Published: (2023)
Machine Learning: a Lecture Note
by: Cho, Kyunghyun
Published: (2025)
by: Cho, Kyunghyun
Published: (2025)
A Brief Introduction to Causal Inference in Machine Learning
by: Cho, Kyunghyun
Published: (2024)
by: Cho, Kyunghyun
Published: (2024)
Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications
by: Halfon, Alon, et al.
Published: (2024)
by: Halfon, Alon, et al.
Published: (2024)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
by: Wang, Atticus, et al.
Published: (2025)
by: Wang, Atticus, et al.
Published: (2025)
Simple Mechanisms for Representing, Indexing and Manipulating Concepts
by: Li, Yuanzhi, et al.
Published: (2023)
by: Li, Yuanzhi, et al.
Published: (2023)
Large Language Model Enhanced Particle Swarm Optimization for Hyperparameter Tuning for Deep Learning Models
by: Hameed, Saad, et al.
Published: (2025)
by: Hameed, Saad, et al.
Published: (2025)
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
by: Xie, Tong, et al.
Published: (2025)
by: Xie, Tong, et al.
Published: (2025)
It's Not That Simple. An Analysis of Simple Test-Time Scaling
by: Wu, Guojun
Published: (2025)
by: Wu, Guojun
Published: (2025)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
by: Jaiswal, Ajay, et al.
Published: (2023)
by: Jaiswal, Ajay, et al.
Published: (2023)
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
by: Martinez, Matias
Published: (2024)
by: Martinez, Matias
Published: (2024)
Olica: Efficient Structured Pruning of Large Language Models without Retraining
by: He, Jiujun, et al.
Published: (2025)
by: He, Jiujun, et al.
Published: (2025)
Similar Items
-
Show Your Work with Confidence: Confidence Bands for Tuning Curves
by: Lourie, Nicholas, et al.
Published: (2023) -
Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
by: Lourie, Nicholas, et al.
Published: (2025) -
Neural Neural Scaling Laws
by: Hu, Michael Y., et al.
Published: (2026) -
Aioli: A Unified Optimization Framework for Language Model Data Mixing
by: Chen, Mayee F., et al.
Published: (2024) -
Hyperparameters in Continual Learning: A Reality Check
by: Cha, Sungmin, et al.
Published: (2024)