Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Hariri, Mohsen, Samandar, Amirhossein, Hinczewski, Michael, Chaudhary, Vipin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scorio.jl: A Julia package for ranking stochastic responses
by: Hariri, Mohsen, et al.
Published: (2026)
by: Hariri, Mohsen, et al.
Published: (2026)
Ranking Reasoning LLMs under Test-Time Scaling
by: Hariri, Mohsen, et al.
Published: (2026)
by: Hariri, Mohsen, et al.
Published: (2026)
A Statistical Hypothesis Testing Framework for Data Misappropriation Detection in Large Language Models
by: Cai, Yinpeng, et al.
Published: (2025)
by: Cai, Yinpeng, et al.
Published: (2025)
Foundations of Top-$k$ Decoding For Language Models
by: Noarov, Georgy, et al.
Published: (2025)
by: Noarov, Georgy, et al.
Published: (2025)
Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference
by: Golowich, Noah, et al.
Published: (2026)
by: Golowich, Noah, et al.
Published: (2026)
Large Language Models Must Be Taught to Know What They Don't Know
by: Kapoor, Sanyam, et al.
Published: (2024)
by: Kapoor, Sanyam, et al.
Published: (2024)
RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models
by: Hao, Sai, et al.
Published: (2026)
by: Hao, Sai, et al.
Published: (2026)
Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
by: Foster, Dylan J., et al.
Published: (2025)
by: Foster, Dylan J., et al.
Published: (2025)
Bayesian Mixture-of-Experts: Towards Making LLMs Know What They Don't Know
by: Li, Albus Yizhuo
Published: (2025)
by: Li, Albus Yizhuo
Published: (2025)
Towards Bayesian Data Selection
by: Rodemann, Julian
Published: (2024)
by: Rodemann, Julian
Published: (2024)
A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation
by: Sarmah, Bhaskarjit, et al.
Published: (2024)
by: Sarmah, Bhaskarjit, et al.
Published: (2024)
Counterfactual reasoning: an analysis of in-context emergence
by: Miller, Moritz, et al.
Published: (2025)
by: Miller, Moritz, et al.
Published: (2025)
Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
by: Lin, Licong, et al.
Published: (2023)
by: Lin, Licong, et al.
Published: (2023)
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
The Coverage Principle: How Pre-Training Enables Post-Training
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Reasoning with Sampling: Cutting at Decision Points
by: Zhou, Felix, et al.
Published: (2026)
by: Zhou, Felix, et al.
Published: (2026)
Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods
by: Hu, Xinyang, et al.
Published: (2024)
by: Hu, Xinyang, et al.
Published: (2024)
Phase Transitions in the Output Distribution of Large Language Models
by: Arnold, Julian, et al.
Published: (2024)
by: Arnold, Julian, et al.
Published: (2024)
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025)
by: Chen, Yanda, et al.
Published: (2025)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
by: Svete, Anej, et al.
Published: (2026)
by: Svete, Anej, et al.
Published: (2026)
A Survey on Large Language Model-based Agents for Statistics and Data Science
by: Sun, Maojun, et al.
Published: (2024)
by: Sun, Maojun, et al.
Published: (2024)
MESSY Estimation: Maximum-Entropy based Stochastic and Symbolic densitY Estimation
by: Tohme, Tony, et al.
Published: (2023)
by: Tohme, Tony, et al.
Published: (2023)
LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning
by: Cao, Junyu, et al.
Published: (2026)
by: Cao, Junyu, et al.
Published: (2026)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
by: Chen, Zhipeng, et al.
Published: (2025)
by: Chen, Zhipeng, et al.
Published: (2025)
PassNet: Scaling Large Language Models for Graph Compiler Pass Generation
by: Liu, Yiqun, et al.
Published: (2026)
by: Liu, Yiqun, et al.
Published: (2026)
LLM Cyber Evaluations Don't Capture Real-World Risk
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
by: Leemann, Tobias, et al.
Published: (2024)
by: Leemann, Tobias, et al.
Published: (2024)
L$^2$M: Mutual Information Scaling Law for Long-Context Language Modeling
by: Chen, Zhuo, et al.
Published: (2025)
by: Chen, Zhuo, et al.
Published: (2025)
A Computational Theory for Efficient Mini Agent Evaluation with Causal Guarantees
by: Yan, Hedong
Published: (2025)
by: Yan, Hedong
Published: (2025)
Beyond Demand Estimation: Consumer Surplus Evaluation via Cumulative Propensity Weights
by: Bian, Zeyu, et al.
Published: (2026)
by: Bian, Zeyu, et al.
Published: (2026)
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
by: Khan, Imran
Published: (2025)
by: Khan, Imran
Published: (2025)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
Don't Shoot The Breeze: Topic Continuity Model Using Nonlinear Naive Bayes With Attention
by: Pi, Shu-Ting, et al.
Published: (2026)
by: Pi, Shu-Ting, et al.
Published: (2026)
In-Context Environments Induce Evaluation-Awareness in Language Models
by: Chaudhary, Maheep
Published: (2026)
by: Chaudhary, Maheep
Published: (2026)
Perturbative adaptive importance sampling for Bayesian LOO cross-validation
by: Chang, Joshua C, et al.
Published: (2024)
by: Chang, Joshua C, et al.
Published: (2024)
Retrieval-Augmented Generation as Noisy In-Context Learning: A Unified Theory and Risk Bounds
by: Guo, Yang, et al.
Published: (2025)
by: Guo, Yang, et al.
Published: (2025)
On the Statistical Capacity of Deep Generative Models
by: Tam, Edric, et al.
Published: (2025)
by: Tam, Edric, et al.
Published: (2025)
On the Provable Performance Guarantee of Efficient Reasoning Models
by: Zeng, Hao, et al.
Published: (2025)
by: Zeng, Hao, et al.
Published: (2025)
Counterfactual Generative Modeling with Variational Causal Inference
by: Wu, Yulun, et al.
Published: (2024)
by: Wu, Yulun, et al.
Published: (2024)
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
by: Dong, Zihan, et al.
Published: (2026)
by: Dong, Zihan, et al.
Published: (2026)
Similar Items
-
Scorio.jl: A Julia package for ranking stochastic responses
by: Hariri, Mohsen, et al.
Published: (2026) -
Ranking Reasoning LLMs under Test-Time Scaling
by: Hariri, Mohsen, et al.
Published: (2026) -
A Statistical Hypothesis Testing Framework for Data Misappropriation Detection in Large Language Models
by: Cai, Yinpeng, et al.
Published: (2025) -
Foundations of Top-$k$ Decoding For Language Models
by: Noarov, Georgy, et al.
Published: (2025) -
Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference
by: Golowich, Noah, et al.
Published: (2026)