Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Dekoninck, Jasper, Baader, Maximilian, Vechev, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Generation of Bias-Eliciting Questions for LLMs
by: Staab, Robin, et al.
Published: (2025)
by: Staab, Robin, et al.
Published: (2025)
A Unified Approach to Routing and Cascading for LLMs
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Evading Data Contamination Detection for Language Models is (too) Easy
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
Learning from Saturated Data: Signals Beyond Correctness for LLM Training
by: Hiss, Hanno, et al.
Published: (2026)
by: Hiss, Hanno, et al.
Published: (2026)
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
by: Petrov, Ivo, et al.
Published: (2026)
by: Petrov, Ivo, et al.
Published: (2026)
ConStat: Performance-Based Contamination Detection in Large Language Models
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical Proofs
by: Dekoninck, Jasper, et al.
Published: (2025)
by: Dekoninck, Jasper, et al.
Published: (2025)
AutoBaxBuilder: Bootstrapping Code Security Benchmarking
by: von Arx, Tobias, et al.
Published: (2025)
by: von Arx, Tobias, et al.
Published: (2025)
Ward: Provable RAG Dataset Inference via LLM Watermarks
by: Jovanović, Nikola, et al.
Published: (2024)
by: Jovanović, Nikola, et al.
Published: (2024)
Controlled Text Generation via Language Model Arithmetic
by: Dekoninck, Jasper, et al.
Published: (2023)
by: Dekoninck, Jasper, et al.
Published: (2023)
Constrained Decoding of Diffusion LLMs with Context-Free Grammars
by: Mündler, Niels, et al.
Published: (2025)
by: Mündler, Niels, et al.
Published: (2025)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
by: Mündler, Niels, et al.
Published: (2023)
by: Mündler, Niels, et al.
Published: (2023)
ToolFuzz -- Automated Agent Tool Testing
by: Milev, Ivan, et al.
Published: (2025)
by: Milev, Ivan, et al.
Published: (2025)
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
by: Ding, Wenxuan, et al.
Published: (2026)
by: Ding, Wenxuan, et al.
Published: (2026)
No LLM is Free From Bias: A Comprehensive Study of Bias Evaluation in Large Language Models
by: Kumar, Charaka Vinayak, et al.
Published: (2025)
by: Kumar, Charaka Vinayak, et al.
Published: (2025)
Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning
by: Wu, Xuyang, et al.
Published: (2025)
by: Wu, Xuyang, et al.
Published: (2025)
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026)
by: Hwang, Yerin, et al.
Published: (2026)
Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets
by: Yukhymenko, Hanna, et al.
Published: (2026)
by: Yukhymenko, Hanna, et al.
Published: (2026)
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR
by: Egashira, Kazuki, et al.
Published: (2026)
by: Egashira, Kazuki, et al.
Published: (2026)
CascadeDebate: Multi-Agent Deliberation for Cost-Aware LLM Cascades
by: Chang, Raeyoung, et al.
Published: (2026)
by: Chang, Raeyoung, et al.
Published: (2026)
Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems
by: Chhabra, Mukul, et al.
Published: (2026)
by: Chhabra, Mukul, et al.
Published: (2026)
Competition-Level Problems are Effective LLM Evaluators
by: Huang, Yiming, et al.
Published: (2023)
by: Huang, Yiming, et al.
Published: (2023)
Bayesian Calibration of Win Rate Estimation with LLM Evaluators
by: Gao, Yicheng, et al.
Published: (2024)
by: Gao, Yicheng, et al.
Published: (2024)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
by: Dekoninck, Jasper, et al.
Published: (2026)
by: Dekoninck, Jasper, et al.
Published: (2026)
A Synthetic Dataset for Personal Attribute Inference
by: Yukhymenko, Hanna, et al.
Published: (2024)
by: Yukhymenko, Hanna, et al.
Published: (2024)
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation
by: Ghosh, Himel, et al.
Published: (2026)
by: Ghosh, Himel, et al.
Published: (2026)
BaxBench: Can LLMs Generate Correct and Secure Backends?
by: Vero, Mark, et al.
Published: (2025)
by: Vero, Mark, et al.
Published: (2025)
LLM-Guided Synthetic Augmentation (LGSA) for Mitigating Bias in AI Systems
by: Karri, Sai Suhruth Reddy, et al.
Published: (2025)
by: Karri, Sai Suhruth Reddy, et al.
Published: (2025)
MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings
by: Zhang, Yiqun, et al.
Published: (2026)
by: Zhang, Yiqun, et al.
Published: (2026)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
by: Lai, Peng, et al.
Published: (2026)
by: Lai, Peng, et al.
Published: (2026)
Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings
by: Hackl, Veronika, et al.
Published: (2023)
by: Hackl, Veronika, et al.
Published: (2023)
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
by: Liu, Jiayu, et al.
Published: (2025)
by: Liu, Jiayu, et al.
Published: (2025)
Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
by: Balamurali, Sai Shridhar, et al.
Published: (2025)
by: Balamurali, Sai Shridhar, et al.
Published: (2025)
DA-Cramming: Enhancing Cost-Effective Language Model Pretraining with Dependency Agreement Integration
by: Kuo, Martin, et al.
Published: (2023)
by: Kuo, Martin, et al.
Published: (2023)
Improving Data Efficiency via Curating LLM-Driven Rating Systems
by: Pang, Jinlong, et al.
Published: (2024)
by: Pang, Jinlong, et al.
Published: (2024)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
by: Ding, Dujian, et al.
Published: (2024)
by: Ding, Dujian, et al.
Published: (2024)
Similar Items
-
Adaptive Generation of Bias-Eliciting Questions for LLMs
by: Staab, Robin, et al.
Published: (2025) -
A Unified Approach to Routing and Cascading for LLMs
by: Dekoninck, Jasper, et al.
Published: (2024) -
Evading Data Contamination Detection for Language Models is (too) Easy
by: Dekoninck, Jasper, et al.
Published: (2024) -
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025) -
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
by: Balunović, Mislav, et al.
Published: (2025)