LLM-as-Judge on a Budget
Fuente:
arXiv
Salvato in:
| Autori principali: | Saha, Aadirupa, Wagde, Aniket, Kveton, Branislav |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning from a single labeled face and a stream of unlabeled data
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Online Posterior Sampling with a Diffusion Prior
di: Kveton, Branislav, et al.
Pubblicazione: (2024)
di: Kveton, Branislav, et al.
Pubblicazione: (2024)
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
DP-Dueling: Learning from Preference Feedback without Compromising User Privacy
di: Saha, Aadirupa, et al.
Pubblicazione: (2024)
di: Saha, Aadirupa, et al.
Pubblicazione: (2024)
Stop Relying on No-Choice and Do not Repeat the Moves: Optimal, Efficient and Practical Algorithms for Assortment Optimization
di: Saha, Aadirupa, et al.
Pubblicazione: (2024)
di: Saha, Aadirupa, et al.
Pubblicazione: (2024)
Cross-Validated Off-Policy Evaluation
di: Cief, Matej, et al.
Pubblicazione: (2024)
di: Cief, Matej, et al.
Pubblicazione: (2024)
Efficient and Interpretable Bandit Algorithms
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2023)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2023)
Quantitative LLM Judges
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
One Good Source is All You Need: Near-Optimal Regret for Bandits under Heterogeneous Noise
di: Bhat, Amith, et al.
Pubblicazione: (2026)
di: Bhat, Amith, et al.
Pubblicazione: (2026)
Tracking the Best Expert Privately
di: Saha, Aadirupa, et al.
Pubblicazione: (2025)
di: Saha, Aadirupa, et al.
Pubblicazione: (2025)
Experimental Design for Active Transductive Inference in Large Language Models
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
Optimal Design for Human Preference Elicitation
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
Pessimistic Off-Policy Optimization for Learning to Rank
di: Cief, Matej, et al.
Pubblicazione: (2022)
di: Cief, Matej, et al.
Pubblicazione: (2022)
Semi-supervised learning with max-margin graph cuts
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Spectral bandits for smooth graph functions
di: Valko, Michal, et al.
Pubblicazione: (2026)
di: Valko, Michal, et al.
Pubblicazione: (2026)
Online semi-supervised perception: Real-time learning without explicit feedback
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Spectral bandits for smooth graph functions with applications in recommender systems
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
Evidence-based anomaly detection in clinical domains
di: Hauskrecht, Milos, et al.
Pubblicazione: (2026)
di: Hauskrecht, Milos, et al.
Pubblicazione: (2026)
Off-Policy Evaluation from Logged Human Feedback
di: Bhargava, Aniruddha, et al.
Pubblicazione: (2024)
di: Bhargava, Aniruddha, et al.
Pubblicazione: (2024)
Finite-Time Logarithmic Bayes Regret Upper Bounds
di: Atsidakou, Alexia, et al.
Pubblicazione: (2023)
di: Atsidakou, Alexia, et al.
Pubblicazione: (2023)
Partial Policy Gradients for RL in LLMs
di: Mathur, Puneet, et al.
Pubblicazione: (2026)
di: Mathur, Puneet, et al.
Pubblicazione: (2026)
Conditional anomaly detection with soft harmonic functions
di: Valko, Michal, et al.
Pubblicazione: (2026)
di: Valko, Michal, et al.
Pubblicazione: (2026)
Conditional anomaly detection using soft harmonic functions: An application to clinical alerting
di: Valko, Michal, et al.
Pubblicazione: (2026)
di: Valko, Michal, et al.
Pubblicazione: (2026)
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
di: Bose, Avinandan, et al.
Pubblicazione: (2024)
di: Bose, Avinandan, et al.
Pubblicazione: (2024)
Spectral bandits
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
di: Kocák, Tomáš, et al.
Pubblicazione: (2026)
ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
di: Chittepu, Yaswanth, et al.
Pubblicazione: (2025)
di: Chittepu, Yaswanth, et al.
Pubblicazione: (2025)
Strategic Linear Contextual Bandits
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2024)
di: Buening, Thomas Kleine, et al.
Pubblicazione: (2024)
Instance-Optimal Estimation with Multiple LLM Judges on a Budget
di: Lee, Junghyun, et al.
Pubblicazione: (2026)
di: Lee, Junghyun, et al.
Pubblicazione: (2026)
Agentic Planning with Reasoning for Image Styling via Offline RL
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2026)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2026)
An Efficient Plugin Method for Metric Optimization of Black-Box Models
di: Devic, Siddartha, et al.
Pubblicazione: (2025)
di: Devic, Siddartha, et al.
Pubblicazione: (2025)
On the Vulnerability of Fairness Constrained Learning to Malicious Noise
di: Blum, Avrim, et al.
Pubblicazione: (2023)
di: Blum, Avrim, et al.
Pubblicazione: (2023)
RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
di: Fernandez, Nigel, et al.
Pubblicazione: (2025)
di: Fernandez, Nigel, et al.
Pubblicazione: (2025)
Learning to Allocate Resources with Censored Feedback
di: Montanari, Giovanni, et al.
Pubblicazione: (2026)
di: Montanari, Giovanni, et al.
Pubblicazione: (2026)
FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain
di: Deb, Rohan, et al.
Pubblicazione: (2025)
di: Deb, Rohan, et al.
Pubblicazione: (2025)
Selective Uncertainty Propagation in Offline RL
di: Krishnamurthy, Sanath Kumar, et al.
Pubblicazione: (2023)
di: Krishnamurthy, Sanath Kumar, et al.
Pubblicazione: (2023)
Active Learning for Direct Preference Optimization
di: Kveton, Branislav, et al.
Pubblicazione: (2025)
di: Kveton, Branislav, et al.
Pubblicazione: (2025)
Language-Model Prior Overcomes Cold-Start Items
di: Wang, Shiyu, et al.
Pubblicazione: (2024)
di: Wang, Shiyu, et al.
Pubblicazione: (2024)
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
di: Thekumparampil, Kiran Koshy, et al.
Pubblicazione: (2024)
di: Thekumparampil, Kiran Koshy, et al.
Pubblicazione: (2024)
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
di: Ozkara, Kaan, et al.
Pubblicazione: (2024)
di: Ozkara, Kaan, et al.
Pubblicazione: (2024)
AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Learning from a single labeled face and a stream of unlabeled data
di: Kveton, Branislav, et al.
Pubblicazione: (2026) -
Online Posterior Sampling with a Diffusion Prior
di: Kveton, Branislav, et al.
Pubblicazione: (2024) -
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024) -
DP-Dueling: Learning from Preference Feedback without Compromising User Privacy
di: Saha, Aadirupa, et al.
Pubblicazione: (2024) -
Stop Relying on No-Choice and Do not Repeat the Moves: Optimal, Efficient and Practical Algorithms for Assortment Optimization
di: Saha, Aadirupa, et al.
Pubblicazione: (2024)