Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
Fuente:
arXiv
Saved in:
| Main Authors: | Bowyer, Sam, Aitchison, Laurence, Ivanova, Desi R. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
xAI-Drop: Don't Use What You Cannot Explain
by: De Luca, Vincenzo Marco, et al.
Published: (2024)
by: De Luca, Vincenzo Marco, et al.
Published: (2024)
Extending Epistemic Uncertainty Beyond Parameters Would Assist in Designing Reliable LLMs
by: Nguyen-Hien, T. Duy, et al.
Published: (2025)
by: Nguyen-Hien, T. Duy, et al.
Published: (2025)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
by: Zhang, Jiefu, et al.
Published: (2026)
by: Zhang, Jiefu, et al.
Published: (2026)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target
by: Kim, Taesan, et al.
Published: (2026)
by: Kim, Taesan, et al.
Published: (2026)
Coordinate Heart System: A Geometric Framework for Emotion Representation
by: Al-Desi, Omar
Published: (2025)
by: Al-Desi, Omar
Published: (2025)
Adam-mini: Use Fewer Learning Rates To Gain More
by: Zhang, Yushun, et al.
Published: (2024)
by: Zhang, Yushun, et al.
Published: (2024)
Transformers Don't In-Context Learn Least Squares Regression
by: Hill, Joshua, et al.
Published: (2025)
by: Hill, Joshua, et al.
Published: (2025)
JADAI: Jointly Amortizing Adaptive Design and Bayesian Inference
by: Bracher, Niels, et al.
Published: (2025)
by: Bracher, Niels, et al.
Published: (2025)
Leveraging Self-Consistency for Data-Efficient Amortized Bayesian Inference
by: Schmitt, Marvin, et al.
Published: (2023)
by: Schmitt, Marvin, et al.
Published: (2023)
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
by: Bowyer, Sam, et al.
Published: (2026)
by: Bowyer, Sam, et al.
Published: (2026)
Benchmarking is Broken -- Don't Let AI be its Own Judge
by: Cheng, Zerui, et al.
Published: (2025)
by: Cheng, Zerui, et al.
Published: (2025)
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
by: Khoriaty, Matthew, et al.
Published: (2025)
by: Khoriaty, Matthew, et al.
Published: (2025)
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
by: Kim, Hyunseung, et al.
Published: (2024)
by: Kim, Hyunseung, et al.
Published: (2024)
Don't Waste Your Time: Early Stopping Cross-Validation
by: Bergman, Edward, et al.
Published: (2024)
by: Bergman, Edward, et al.
Published: (2024)
LLM Cyber Evaluations Don't Capture Real-World Risk
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
Don't Retrain, Align: Adapting Autoregressive LMs to Diffusion LMs via Representation Alignment
by: Peng, Fred Zhangzhi, et al.
Published: (2026)
by: Peng, Fred Zhangzhi, et al.
Published: (2026)
Don't Lag, RAG: Training-Free Adversarial Detection Using RAG
by: Kazoom, Roie, et al.
Published: (2025)
by: Kazoom, Roie, et al.
Published: (2025)
Don't be lazy: CompleteP enables compute-efficient deep transformers
by: Dey, Nolan, et al.
Published: (2025)
by: Dey, Nolan, et al.
Published: (2025)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025)
by: Farnik, Lucy, et al.
Published: (2025)
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025)
by: Chen, Yanda, et al.
Published: (2025)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
by: Park, Young-Jin, et al.
Published: (2025)
by: Park, Young-Jin, et al.
Published: (2025)
Know What You Don't Know: Selective Prediction for Early Exit DNNs
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
Don't throw the baby out with the bathwater: How and why deep learning for ARC
by: Cole, Jack, et al.
Published: (2025)
by: Cole, Jack, et al.
Published: (2025)
Trust, or Don't Predict: Introducing the CWSA Family for Confidence-Aware Model Evaluation
by: Shahnazari, Kourosh, et al.
Published: (2025)
by: Shahnazari, Kourosh, et al.
Published: (2025)
Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)
by: Lade, Ankit Hemant, et al.
Published: (2026)
by: Lade, Ankit Hemant, et al.
Published: (2026)
mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
by: Yang, Yongyi, et al.
Published: (2026)
by: Yang, Yongyi, et al.
Published: (2026)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Don't stop me now: Rethinking Validation Criteria for Model Parameter Selection
by: Apicella, Andrea, et al.
Published: (2026)
by: Apicella, Andrea, et al.
Published: (2026)
Nonasymptotic CLT and Error Bounds for Two-Time-Scale Stochastic Approximation
by: Kong, Seo Taek, et al.
Published: (2025)
by: Kong, Seo Taek, et al.
Published: (2025)
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
What LLMs Think When You Don't Tell Them What to Think About?
by: Kwon, Yongchan, et al.
Published: (2026)
by: Kwon, Yongchan, et al.
Published: (2026)
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
by: Sokar, Ghada, et al.
Published: (2024)
by: Sokar, Ghada, et al.
Published: (2024)
Don't Push the Button! Exploring Data Leakage Risks in Machine Learning and Transfer Learning
by: Apicella, Andrea, et al.
Published: (2024)
by: Apicella, Andrea, et al.
Published: (2024)
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
by: Khan, Imran
Published: (2025)
by: Khan, Imran
Published: (2025)
Trust, Don't Trust, or Flip: Robust Preference-Based Reinforcement Learning with Multi-Expert Feedback
by: Hosseini, Seyed Amir, et al.
Published: (2026)
by: Hosseini, Seyed Amir, et al.
Published: (2026)
Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning
by: Poole, Benjamin, et al.
Published: (2026)
by: Poole, Benjamin, et al.
Published: (2026)
Similar Items
-
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025) -
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024) -
xAI-Drop: Don't Use What You Cannot Explain
by: De Luca, Vincenzo Marco, et al.
Published: (2024) -
Extending Epistemic Uncertainty Beyond Parameters Would Assist in Designing Reliable LLMs
by: Nguyen-Hien, T. Duy, et al.
Published: (2025) -
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
by: Zhang, Jiefu, et al.
Published: (2026)