Model Consistency as a Cheap yet Predictive Proxy for LLM Elo Scores
Fuente:
arXiv
Guardado en:
| Autores principales: | Ramaswamy, Ashwin, Demeure, Nestor, Rrapaj, Ermal |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Adapting Reinforcement Learning Agents to New Tasks: Insights from Q-Values
por: Ramaswamy, Ashwin, et al.
Publicado: (2024)
por: Ramaswamy, Ashwin, et al.
Publicado: (2024)
Modeling LLM Agent Reviewer Dynamics in Elo-Ranked Review System
por: Huang, Hsiang-Wei, et al.
Publicado: (2026)
por: Huang, Hsiang-Wei, et al.
Publicado: (2026)
Exact block encoding of imaginary time evolution with universal quantum neural networks
por: Rrapaj, Ermal, et al.
Publicado: (2024)
por: Rrapaj, Ermal, et al.
Publicado: (2024)
Predicting LLM Reasoning Performance with Small Proxy Model
por: Koh, Woosung, et al.
Publicado: (2025)
por: Koh, Woosung, et al.
Publicado: (2025)
Tournament of Prompts: Evolving LLM Instructions Through Structured Debates and Elo Ratings
por: Nair, Anirudh, et al.
Publicado: (2025)
por: Nair, Anirudh, et al.
Publicado: (2025)
Cheap Talking Algorithms
por: Condorelli, Daniele, et al.
Publicado: (2023)
por: Condorelli, Daniele, et al.
Publicado: (2023)
Is Elo Rating Reliable? A Study Under Model Misspecification
por: Tang, Shange, et al.
Publicado: (2025)
por: Tang, Shange, et al.
Publicado: (2025)
ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference
por: Li, Junjie, et al.
Publicado: (2026)
por: Li, Junjie, et al.
Publicado: (2026)
Elo-Evolve: A Co-evolutionary Framework for Language Model Alignment
por: Zhao, Jing, et al.
Publicado: (2026)
por: Zhao, Jing, et al.
Publicado: (2026)
Cycle-Consistent Search: Question Reconstructability as a Proxy Reward for Search Agent Training
por: An, Sohyun, et al.
Publicado: (2026)
por: An, Sohyun, et al.
Publicado: (2026)
AGI-Elo: How Far Are We From Mastering A Task?
por: Sun, Shuo, et al.
Publicado: (2025)
por: Sun, Shuo, et al.
Publicado: (2025)
Classical optimization with imaginary time block encoding on quantum computers: The MaxCut problem
por: Zhong, Dawei, et al.
Publicado: (2024)
por: Zhong, Dawei, et al.
Publicado: (2024)
Fast Proxies for LLM Robustness Evaluation
por: Beyer, Tim, et al.
Publicado: (2025)
por: Beyer, Tim, et al.
Publicado: (2025)
DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs
por: Huang, Minghui
Publicado: (2025)
por: Huang, Minghui
Publicado: (2025)
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
por: Wang, Guanghui, et al.
Publicado: (2025)
por: Wang, Guanghui, et al.
Publicado: (2025)
The Impact of LLM Self-Consistency and Reasoning Effort on Automated Scoring Accuracy and Cost
por: Frohn, Scott
Publicado: (2026)
por: Frohn, Scott
Publicado: (2026)
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
por: Li, Zongxia, et al.
Publicado: (2024)
por: Li, Zongxia, et al.
Publicado: (2024)
Advancing Harmful Content Detection in Organizational Research: Integrating Large Language Models with Elo Rating System
por: Akben, Mustafa, et al.
Publicado: (2025)
por: Akben, Mustafa, et al.
Publicado: (2025)
Detecting LLM-assisted writing in scientific communication: Are we there yet?
por: Lazebnik, Teddy, et al.
Publicado: (2024)
por: Lazebnik, Teddy, et al.
Publicado: (2024)
Type-Less yet Type-Aware Inductive Link Prediction with Pretrained Language Models
por: De Bellis, Alessandro, et al.
Publicado: (2025)
por: De Bellis, Alessandro, et al.
Publicado: (2025)
OpenSeal: Good, Fast, and Cheap Construction of an Open-Source Southeast Asian LLM via Parallel Data
por: Nguyen, Tan Sang, et al.
Publicado: (2026)
por: Nguyen, Tan Sang, et al.
Publicado: (2026)
Peak + Accumulation: A Proxy-Level Scoring Formula for Multi-Turn LLM Attack Detection
por: Corll, J Alex
Publicado: (2026)
por: Corll, J Alex
Publicado: (2026)
PredictaBoard: Benchmarking LLM Score Predictability
por: Pacchiardi, Lorenzo, et al.
Publicado: (2025)
por: Pacchiardi, Lorenzo, et al.
Publicado: (2025)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
por: Kiyani, Shayan, et al.
Publicado: (2026)
por: Kiyani, Shayan, et al.
Publicado: (2026)
Uncovering Latent Bias in LLM-Based Emergency Department Triage Through Proxy Variables
por: Zhang, Ethan
Publicado: (2026)
por: Zhang, Ethan
Publicado: (2026)
Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble
por: Wang, Bingchen, et al.
Publicado: (2025)
por: Wang, Bingchen, et al.
Publicado: (2025)
MCP Bridge: A Lightweight, LLM-Agnostic RESTful Proxy for Model Context Protocol Servers
por: Ahmadi, Arash, et al.
Publicado: (2025)
por: Ahmadi, Arash, et al.
Publicado: (2025)
A Proxy Consistency Loss for Grounded Fusion of Earth Observation and Location Encoders
por: Wang, Zhongying, et al.
Publicado: (2026)
por: Wang, Zhongying, et al.
Publicado: (2026)
ProxyWar: Dynamic Assessment of LLM Code Generation in Game Arenas
por: Peng, Wenjun, et al.
Publicado: (2026)
por: Peng, Wenjun, et al.
Publicado: (2026)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
por: Zhu, Yu, et al.
Publicado: (2024)
por: Zhu, Yu, et al.
Publicado: (2024)
When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search
por: Robertson, John T., et al.
Publicado: (2026)
por: Robertson, John T., et al.
Publicado: (2026)
Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation
por: Sun, Jiashuo, et al.
Publicado: (2026)
por: Sun, Jiashuo, et al.
Publicado: (2026)
Decoding Generalization from Memorization in Deep Neural Networks
por: Ketha, Simran, et al.
Publicado: (2025)
por: Ketha, Simran, et al.
Publicado: (2025)
Language Complexity Measurement as a Noisy Zero-Shot Proxy for Evaluating LLM Performance
por: Moell, Birger, et al.
Publicado: (2025)
por: Moell, Birger, et al.
Publicado: (2025)
AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
por: Guo, Taicheng, et al.
Publicado: (2026)
por: Guo, Taicheng, et al.
Publicado: (2026)
Rich Insights from Cheap Signals: Efficient Evaluations via Tensor Factorization
por: Polo, Felipe Maia, et al.
Publicado: (2026)
por: Polo, Felipe Maia, et al.
Publicado: (2026)
Detecting Prefix Bias in LLM-based Reward Models
por: Kumar, Ashwin, et al.
Publicado: (2025)
por: Kumar, Ashwin, et al.
Publicado: (2025)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
por: Kossen, Jannik, et al.
Publicado: (2024)
por: Kossen, Jannik, et al.
Publicado: (2024)
EasyRec: Simple yet Effective Language Models for Recommendation
por: Ren, Xubin, et al.
Publicado: (2024)
por: Ren, Xubin, et al.
Publicado: (2024)
GCCM: Enhancing Generative Graph Prediction via Contrastive Consistency Model
por: Ma, Shaozhen, et al.
Publicado: (2026)
por: Ma, Shaozhen, et al.
Publicado: (2026)
Ejemplares similares
-
Towards Adapting Reinforcement Learning Agents to New Tasks: Insights from Q-Values
por: Ramaswamy, Ashwin, et al.
Publicado: (2024) -
Modeling LLM Agent Reviewer Dynamics in Elo-Ranked Review System
por: Huang, Hsiang-Wei, et al.
Publicado: (2026) -
Exact block encoding of imaginary time evolution with universal quantum neural networks
por: Rrapaj, Ermal, et al.
Publicado: (2024) -
Predicting LLM Reasoning Performance with Small Proxy Model
por: Koh, Woosung, et al.
Publicado: (2025) -
Tournament of Prompts: Evolving LLM Instructions Through Structured Debates and Elo Ratings
por: Nair, Anirudh, et al.
Publicado: (2025)