The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Amin, Adil |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling
par: Amin, Adil
Publié: (2026)
par: Amin, Adil
Publié: (2026)
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
par: Alzahrani, Norah, et autres
Publié: (2024)
par: Alzahrani, Norah, et autres
Publié: (2024)
The Leaderboard Illusion
par: Singh, Shivalika, et autres
Publié: (2025)
par: Singh, Shivalika, et autres
Publié: (2025)
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard
par: Rao, Varun, et autres
Publié: (2025)
par: Rao, Varun, et autres
Publié: (2025)
When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation
par: Faisal, Faizan
Publié: (2026)
par: Faisal, Faizan
Publié: (2026)
Early Stopping for Large Reasoning Models via Confidence Dynamics
par: Hosseini, Parsa, et autres
Publié: (2026)
par: Hosseini, Parsa, et autres
Publié: (2026)
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard
par: Topsakal, Oguzhan, et autres
Publié: (2024)
par: Topsakal, Oguzhan, et autres
Publié: (2024)
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
par: Chen, Wenting, et autres
Publié: (2025)
par: Chen, Wenting, et autres
Publié: (2025)
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
par: Alqahtani, Sawsan, et autres
Publié: (2026)
par: Alqahtani, Sawsan, et autres
Publié: (2026)
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
par: Fronsdal, Kai, et autres
Publié: (2024)
par: Fronsdal, Kai, et autres
Publié: (2024)
NextLocLLM: Location Semantics Modeling and Coordinate-Based Next Location Prediction with LLMs
par: Liu, Shuai, et autres
Publié: (2024)
par: Liu, Shuai, et autres
Publié: (2024)
Do We Need Frontier Models to Verify Mathematical Proofs?
par: Naik, Aaditya, et autres
Publié: (2026)
par: Naik, Aaditya, et autres
Publié: (2026)
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models
par: Saha, Partha Pratim, et autres
Publié: (2026)
par: Saha, Partha Pratim, et autres
Publié: (2026)
Learning to Reason at the Frontier of Learnability
par: Foster, Thomas, et autres
Publié: (2025)
par: Foster, Thomas, et autres
Publié: (2025)
Towards the Next Frontier in Speech Representation Learning Using Disentanglement
par: Krishna, Varun, et autres
Publié: (2024)
par: Krishna, Varun, et autres
Publié: (2024)
Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation
par: Glocker, Kevin, et autres
Publié: (2025)
par: Glocker, Kevin, et autres
Publié: (2025)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
par: Kang, Feiyang, et autres
Publié: (2025)
par: Kang, Feiyang, et autres
Publié: (2025)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
par: Liu, Zihan, et autres
Publié: (2024)
par: Liu, Zihan, et autres
Publié: (2024)
Explaining Large Language Models with gSMILE
par: Dehghani, Zeinab, et autres
Publié: (2025)
par: Dehghani, Zeinab, et autres
Publié: (2025)
Examining Gender and Power on Wikipedia Through Face and Politeness
par: Soubki, Adil, et autres
Publié: (2024)
par: Soubki, Adil, et autres
Publié: (2024)
A Law of Next-Token Prediction in Large Language Models
par: He, Hangfeng, et autres
Publié: (2024)
par: He, Hangfeng, et autres
Publié: (2024)
Text Knows What, Tables Know When: Clinical Timeline Reconstruction via Retrieval-Augmented Multimodal Alignment
par: Kumar, Sayantan, et autres
Publié: (2026)
par: Kumar, Sayantan, et autres
Publié: (2026)
From Growing to Looping: A Unified View of Iterative Computation in LLMs
par: Kapl, Ferdinand, et autres
Publié: (2026)
par: Kapl, Ferdinand, et autres
Publié: (2026)
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
par: Nagle, Alliot, et autres
Publié: (2026)
par: Nagle, Alliot, et autres
Publié: (2026)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
par: Schaeffer, Rylan, et autres
Publié: (2024)
par: Schaeffer, Rylan, et autres
Publié: (2024)
When Bad Data Leads to Good Models
par: Li, Kenneth, et autres
Publié: (2025)
par: Li, Kenneth, et autres
Publié: (2025)
Stop Listening to Me! How Multi-turn Conversations Can Degrade LLM Reliability
par: Guo, Kevin H., et autres
Publié: (2026)
par: Guo, Kevin H., et autres
Publié: (2026)
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
par: Hu, Tiancheng, et autres
Publié: (2025)
par: Hu, Tiancheng, et autres
Publié: (2025)
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
par: Verdini, Francesco, et autres
Publié: (2024)
par: Verdini, Francesco, et autres
Publié: (2024)
Cautious Next Token Prediction
par: Wang, Yizhou, et autres
Publié: (2025)
par: Wang, Yizhou, et autres
Publié: (2025)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
par: Malek, Alan, et autres
Publié: (2025)
par: Malek, Alan, et autres
Publié: (2025)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
par: Marconato, Emanuele, et autres
Publié: (2024)
par: Marconato, Emanuele, et autres
Publié: (2024)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
par: Shao, Chenze, et autres
Publié: (2024)
par: Shao, Chenze, et autres
Publié: (2024)
What Matters for Model Merging at Scale?
par: Yadav, Prateek, et autres
Publié: (2024)
par: Yadav, Prateek, et autres
Publié: (2024)
When Incentives Backfire, Data Stops Being Human
par: Santy, Sebastin, et autres
Publié: (2025)
par: Santy, Sebastin, et autres
Publié: (2025)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
par: Ballon, Marthe, et autres
Publié: (2026)
par: Ballon, Marthe, et autres
Publié: (2026)
When Attention Sink Emerges in Language Models: An Empirical View
par: Gu, Xiangming, et autres
Publié: (2024)
par: Gu, Xiangming, et autres
Publié: (2024)
AdaptThink: Reasoning Models Can Learn When to Think
par: Zhang, Jiajie, et autres
Publié: (2025)
par: Zhang, Jiajie, et autres
Publié: (2025)
Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL
par: Zheng, Kunhao, et autres
Publié: (2026)
par: Zheng, Kunhao, et autres
Publié: (2026)
When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
par: Xu, Haoming, et autres
Publié: (2026)
par: Xu, Haoming, et autres
Publié: (2026)
Documents similaires
-
Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling
par: Amin, Adil
Publié: (2026) -
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
par: Alzahrani, Norah, et autres
Publié: (2024) -
The Leaderboard Illusion
par: Singh, Shivalika, et autres
Publié: (2025) -
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard
par: Rao, Varun, et autres
Publié: (2025) -
When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation
par: Faisal, Faizan
Publié: (2026)