Ranking Reasoning LLMs under Test-Time Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Hariri, Mohsen, Hinczewski, Michael, Ma, Jing, Chaudhary, Vipin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scorio.jl: A Julia package for ranking stochastic responses
by: Hariri, Mohsen, et al.
Published: (2026)
by: Hariri, Mohsen, et al.
Published: (2026)
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
by: Hariri, Mohsen, et al.
Published: (2025)
by: Hariri, Mohsen, et al.
Published: (2025)
Thermodynamic Performance Limits for Score-Based Diffusion Models
by: Kodama, Nathan X., et al.
Published: (2025)
by: Kodama, Nathan X., et al.
Published: (2025)
On Ranking-based Tests of Independence
by: Limnios, Myrto, et al.
Published: (2024)
by: Limnios, Myrto, et al.
Published: (2024)
Max-Rank: Efficient Multiple Testing for Conformal Prediction
by: Timans, Alexander, et al.
Published: (2023)
by: Timans, Alexander, et al.
Published: (2023)
Mean Testing under Truncation beyond Gaussian
by: Wang, Yuhao, et al.
Published: (2026)
by: Wang, Yuhao, et al.
Published: (2026)
Compress Then Test: Powerful Kernel Testing in Near-linear Time
by: Domingo-Enrich, Carles, et al.
Published: (2023)
by: Domingo-Enrich, Carles, et al.
Published: (2023)
Dynamic Ranking and Translation Synchronization
by: Araya, Ernesto, et al.
Published: (2022)
by: Araya, Ernesto, et al.
Published: (2022)
On Stopping Times of Power-one Sequential Tests: Tight Lower and Upper Bounds
by: Agrawal, Shubhada, et al.
Published: (2025)
by: Agrawal, Shubhada, et al.
Published: (2025)
The Sample Complexity of Distributed Simple Binary Hypothesis Testing under Information Constraints
by: Kazemi, Hadi, et al.
Published: (2025)
by: Kazemi, Hadi, et al.
Published: (2025)
Entrywise Error Bounds for Spectral Ranking with Semi-Random Adversaries
by: Lee, Dongmin, et al.
Published: (2026)
by: Lee, Dongmin, et al.
Published: (2026)
Parametric Scaling Law of Tuning Bias in Conformal Prediction
by: Zeng, Hao, et al.
Published: (2025)
by: Zeng, Hao, et al.
Published: (2025)
Cyclic Counterfactuals under Shift-Scale Interventions
by: Saha, Saptarshi, et al.
Published: (2025)
by: Saha, Saptarshi, et al.
Published: (2025)
Multiple Testing of Linear Forms for Noisy Matrix Completion
by: Ma, Wanteng, et al.
Published: (2023)
by: Ma, Wanteng, et al.
Published: (2023)
Information-Theoretic Guarantees for Recovering Low-Rank Tensors from Symmetric Rank-One Measurements
by: Kızıldağ, Eren C.
Published: (2025)
by: Kızıldağ, Eren C.
Published: (2025)
Learning from Biased and Costly Data Sources: Minimax-optimal Data Collection under a Budget
by: Harding, Michael O., et al.
Published: (2026)
by: Harding, Michael O., et al.
Published: (2026)
Spacing Test for Fused Lasso
by: Tasaka, Rieko, et al.
Published: (2025)
by: Tasaka, Rieko, et al.
Published: (2025)
Low-Rank Thinning
by: Carrell, Annabelle Michael, et al.
Published: (2025)
by: Carrell, Annabelle Michael, et al.
Published: (2025)
Multi-Armed Sequential Hypothesis Testing by Betting
by: Sandoval, Ricardo J., et al.
Published: (2026)
by: Sandoval, Ricardo J., et al.
Published: (2026)
Sparse Tucker Decomposition and Graph Regularization for High-Dimensional Time Series Forecasting
by: Xia, Sijia, et al.
Published: (2026)
by: Xia, Sijia, et al.
Published: (2026)
Uncertainty Quantification of MLE for Entity Ranking with Covariates
by: Fan, Jianqing, et al.
Published: (2022)
by: Fan, Jianqing, et al.
Published: (2022)
ResCP: Reservoir Conformal Prediction for Time Series Forecasting
by: Neglia, Roberto, et al.
Published: (2025)
by: Neglia, Roberto, et al.
Published: (2025)
Hypothesis Testing for Generalized Thurstone Models
by: Makur, Anuran, et al.
Published: (2025)
by: Makur, Anuran, et al.
Published: (2025)
Policy Testing in Markov Decision Processes
by: Ariu, Kaito, et al.
Published: (2025)
by: Ariu, Kaito, et al.
Published: (2025)
On the Provable Performance Guarantee of Efficient Reasoning Models
by: Zeng, Hao, et al.
Published: (2025)
by: Zeng, Hao, et al.
Published: (2025)
Covariate Assisted Entity Ranking with Sparse Intrinsic Scores
by: Fan, Jianqing, et al.
Published: (2024)
by: Fan, Jianqing, et al.
Published: (2024)
Optimal Differentially Private Ranking from Pairwise Comparisons
by: Cai, T. Tony, et al.
Published: (2025)
by: Cai, T. Tony, et al.
Published: (2025)
Asymptotically Optimal Sequential Testing with Markovian Data
by: Sethi, Alhad, et al.
Published: (2026)
by: Sethi, Alhad, et al.
Published: (2026)
Precise Error Rates for Computationally Efficient Testing
by: Moitra, Ankur, et al.
Published: (2023)
by: Moitra, Ankur, et al.
Published: (2023)
Testing the Feasibility of Linear Programs with Bandit Feedback
by: Gangrade, Aditya, et al.
Published: (2024)
by: Gangrade, Aditya, et al.
Published: (2024)
General Frameworks for Conditional Two-Sample Testing
by: Lee, Seongchan, et al.
Published: (2024)
by: Lee, Seongchan, et al.
Published: (2024)
The Limits of Assumption-free Tests for Algorithm Performance
by: Luo, Yuetian, et al.
Published: (2024)
by: Luo, Yuetian, et al.
Published: (2024)
Kernel Two-Sample Tests for Manifold Data
by: Cheng, Xiuyuan, et al.
Published: (2021)
by: Cheng, Xiuyuan, et al.
Published: (2021)
Scaling Laws are Redundancy Laws
by: Bi, Yuda, et al.
Published: (2025)
by: Bi, Yuda, et al.
Published: (2025)
Statistical Inference under Performativity
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Testing properties of trees in graphical models with covariance queries
by: Burova, Sofiya, et al.
Published: (2026)
by: Burova, Sofiya, et al.
Published: (2026)
Hypothesis Testing over Observable Regimes in Singular Models
by: Plummer, Sean
Published: (2026)
by: Plummer, Sean
Published: (2026)
Statistical Inverse Problems in Hilbert Scales
by: Rastogi, Abhishake
Published: (2022)
by: Rastogi, Abhishake
Published: (2022)
Rectifying Conformity Scores for Better Conditional Coverage
by: Plassier, Vincent, et al.
Published: (2025)
by: Plassier, Vincent, et al.
Published: (2025)
Spectral Ranking Inferences based on General Multiway Comparisons
by: Fan, Jianqing, et al.
Published: (2023)
by: Fan, Jianqing, et al.
Published: (2023)
Similar Items
-
Scorio.jl: A Julia package for ranking stochastic responses
by: Hariri, Mohsen, et al.
Published: (2026) -
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
by: Hariri, Mohsen, et al.
Published: (2025) -
Thermodynamic Performance Limits for Score-Based Diffusion Models
by: Kodama, Nathan X., et al.
Published: (2025) -
On Ranking-based Tests of Independence
by: Limnios, Myrto, et al.
Published: (2024) -
Max-Rank: Efficient Multiple Testing for Conformal Prediction
by: Timans, Alexander, et al.
Published: (2023)