Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Halder, Indranil, Pehlevan, Cengiz |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
by: Halder, Indranil, et al.
Published: (2026)
by: Halder, Indranil, et al.
Published: (2026)
Universal One-third Time Scaling in Learning Peaked Distributions
by: Liu, Yizhou, et al.
Published: (2026)
by: Liu, Yizhou, et al.
Published: (2026)
Error Broadcast and Decorrelation as a Potential Artificial and Natural Learning Mechanism
by: Erdogan, Mete, et al.
Published: (2025)
by: Erdogan, Mete, et al.
Published: (2025)
Score Broadcast and Decorrelation: A General Framework for Broadcast-Based Credit Assignment
by: Uzun, Mustafa, et al.
Published: (2026)
by: Uzun, Mustafa, et al.
Published: (2026)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
by: Lauditi, Clarissa, et al.
Published: (2026)
by: Lauditi, Clarissa, et al.
Published: (2026)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
by: Wang, Yutong, et al.
Published: (2025)
by: Wang, Yutong, et al.
Published: (2025)
Boule or Baguette? A Study on Task Topology, Length Generalization, and the Benefit of Reasoning Traces
by: Tong, William L., et al.
Published: (2026)
by: Tong, William L., et al.
Published: (2026)
SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
by: Zimmer, Max, et al.
Published: (2025)
by: Zimmer, Max, et al.
Published: (2025)
Pixel-Based Similarities as an Alternative to Neural Data for Improving Convolutional Neural Network Adversarial Robustness
by: Attias, Elie, et al.
Published: (2024)
by: Attias, Elie, et al.
Published: (2024)
A Tractable Inference Perspective of Offline RL
by: Liu, Xuejie, et al.
Published: (2023)
by: Liu, Xuejie, et al.
Published: (2023)
Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
by: Dadkhahi, Hamid, et al.
Published: (2025)
by: Dadkhahi, Hamid, et al.
Published: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
Probabilistic Graph Circuits: Deep Generative Models for Tractable Probabilistic Inference over Graphs
by: Papež, Milan, et al.
Published: (2025)
by: Papež, Milan, et al.
Published: (2025)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
by: Huang, Tzu-Heng, et al.
Published: (2025)
by: Huang, Tzu-Heng, et al.
Published: (2025)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
by: Zhu, Xiao, et al.
Published: (2026)
by: Zhu, Xiao, et al.
Published: (2026)
Continuous Mixtures of Tractable Probabilistic Models
by: Correia, Alvaro H. C., et al.
Published: (2022)
by: Correia, Alvaro H. C., et al.
Published: (2022)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
Diversified Scaling Inference in Time Series Foundation Models
by: Hua, Ruijin, et al.
Published: (2026)
by: Hua, Ruijin, et al.
Published: (2026)
Model Collapse Demystified: The Case of Regression
by: Dohmatob, Elvis, et al.
Published: (2024)
by: Dohmatob, Elvis, et al.
Published: (2024)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
by: Li, Xiaochuan, et al.
Published: (2025)
by: Li, Xiaochuan, et al.
Published: (2025)
Don't be lazy: CompleteP enables compute-efficient deep transformers
by: Dey, Nolan, et al.
Published: (2025)
by: Dey, Nolan, et al.
Published: (2025)
Restructuring Tractable Probabilistic Circuits
by: Zhang, Honghua, et al.
Published: (2024)
by: Zhang, Honghua, et al.
Published: (2024)
Demystifying Manifold Constraints in LLM Pre-training
by: An, Kang, et al.
Published: (2026)
by: An, Kang, et al.
Published: (2026)
SymCircuit: Bayesian Structure Inference for Tractable Probabilistic Circuits via Entropy-Regularized Reinforcement Learning
by: Ju, Y. Sungtaek
Published: (2026)
by: Ju, Y. Sungtaek
Published: (2026)
A Random Matrix Theory Perspective on the Consistency of Diffusion Models
by: Wang, Binxu, et al.
Published: (2026)
by: Wang, Binxu, et al.
Published: (2026)
Inference-Time Scaling in Diffusion Models through Iterative Partial Refinement
by: Kang, Taegu, et al.
Published: (2026)
by: Kang, Taegu, et al.
Published: (2026)
Building Expressive and Tractable Probabilistic Generative Models: A Review
by: Sidheekh, Sahil, et al.
Published: (2024)
by: Sidheekh, Sahil, et al.
Published: (2024)
Auto-Prompt Ensemble for LLM Judge
by: Li, Jiajie, et al.
Published: (2025)
by: Li, Jiajie, et al.
Published: (2025)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
Tractable Representation Learning with Probabilistic Circuits
by: Braun, Steven, et al.
Published: (2025)
by: Braun, Steven, et al.
Published: (2025)
Tractable Uncertainty-Aware Meta-Learning
by: Park, Young-Jin, et al.
Published: (2022)
by: Park, Young-Jin, et al.
Published: (2022)
Demystifying MuZero Planning: Interpreting the Learned Model
by: Guei, Hung, et al.
Published: (2024)
by: Guei, Hung, et al.
Published: (2024)
Inference-Time Scaling for Generalist Reward Modeling
by: Liu, Zijun, et al.
Published: (2025)
by: Liu, Zijun, et al.
Published: (2025)
Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
A Dynamical Model of Neural Scaling Laws
by: Bordelon, Blake, et al.
Published: (2024)
by: Bordelon, Blake, et al.
Published: (2024)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
by: Patwari, Rajeev, et al.
Published: (2025)
by: Patwari, Rajeev, et al.
Published: (2025)
Tractable Sharpness-Aware Learning of Probabilistic Circuits
by: Suresh, Hrithik, et al.
Published: (2025)
by: Suresh, Hrithik, et al.
Published: (2025)
On the Tractability of SHAP Explanations under Markovian Distributions
by: Marzouk, Reda, et al.
Published: (2024)
by: Marzouk, Reda, et al.
Published: (2024)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
by: Wagner, Nico, et al.
Published: (2024)
by: Wagner, Nico, et al.
Published: (2024)
REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge
by: Zhang, Yasi, et al.
Published: (2026)
by: Zhang, Yasi, et al.
Published: (2026)
Similar Items
-
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
by: Halder, Indranil, et al.
Published: (2026) -
Universal One-third Time Scaling in Learning Peaked Distributions
by: Liu, Yizhou, et al.
Published: (2026) -
Error Broadcast and Decorrelation as a Potential Artificial and Natural Learning Mechanism
by: Erdogan, Mete, et al.
Published: (2025) -
Score Broadcast and Decorrelation: A General Framework for Broadcast-Based Credit Assignment
by: Uzun, Mustafa, et al.
Published: (2026) -
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
by: Lauditi, Clarissa, et al.
Published: (2026)