Reliable and Efficient Amortized Model-based Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Truong, Sang, Tu, Yuheng, Liang, Percy, Li, Bo, Koyejo, Sanmi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
Extracting books from production language models
by: Ahmed, Ahmed, et al.
Published: (2026)
by: Ahmed, Ahmed, et al.
Published: (2026)
Efficient Prediction of Pass@k Scaling in Large Language Models
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Logits are All We Need to Adapt Closed Models
by: Hiranandani, Gaurush, et al.
Published: (2025)
by: Hiranandani, Gaurush, et al.
Published: (2025)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
by: Zhou, Zhanke, et al.
Published: (2025)
by: Zhou, Zhanke, et al.
Published: (2025)
Domain-Shift-Aware Conformal Prediction for Large Language Models
by: Lin, Zhexiao, et al.
Published: (2025)
by: Lin, Zhexiao, et al.
Published: (2025)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
by: Liu, Ken Ziyu, et al.
Published: (2025)
by: Liu, Ken Ziyu, et al.
Published: (2025)
Uncertainty-Aware Adaptation of Large Language Models for Protein-Protein Interaction Analysis
by: Jantre, Sanket, et al.
Published: (2025)
by: Jantre, Sanket, et al.
Published: (2025)
Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving
by: Kuzmanko, Jonathan
Published: (2025)
by: Kuzmanko, Jonathan
Published: (2025)
Augmented Risk Prediction for the Onset of Alzheimer's Disease from Electronic Health Records with Large Language Models
by: Wang, Jiankun, et al.
Published: (2024)
by: Wang, Jiankun, et al.
Published: (2024)
ICE-ID: A Novel Historical Census Dataset for Longitudinal Identity Resolution
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
by: Hang, Haotian, et al.
Published: (2025)
by: Hang, Haotian, et al.
Published: (2025)
Language Models as Causal Effect Generators
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
Beyond the Hype: Embeddings vs. Prompting for Multiclass Classification Tasks
by: Kokkodis, Marios, et al.
Published: (2025)
by: Kokkodis, Marios, et al.
Published: (2025)
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Unified Representation of Genomic and Biomedical Concepts through Multi-Task, Multi-Source Contrastive Learning
by: Yuan, Hongyi, et al.
Published: (2024)
by: Yuan, Hongyi, et al.
Published: (2024)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
by: Zhou, Cai, et al.
Published: (2026)
by: Zhou, Cai, et al.
Published: (2026)
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Investigating Data Contamination for Pre-training Language Models
by: Jiang, Minhao, et al.
Published: (2024)
by: Jiang, Minhao, et al.
Published: (2024)
Quantifying the Importance of Data Alignment in Downstream Model Performance
by: Chawla, Krrish, et al.
Published: (2025)
by: Chawla, Krrish, et al.
Published: (2025)
UQ: Assessing Language Models on Unsolved Questions
by: Nie, Fan, et al.
Published: (2025)
by: Nie, Fan, et al.
Published: (2025)
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
by: Chen, Edward, et al.
Published: (2025)
by: Chen, Edward, et al.
Published: (2025)
Removing Spurious Correlation from Neural Network Interpretations
by: Fotouhi, Milad, et al.
Published: (2024)
by: Fotouhi, Milad, et al.
Published: (2024)
Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models
by: Nagori, Aditya, et al.
Published: (2025)
by: Nagori, Aditya, et al.
Published: (2025)
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
by: Xu, Yang, et al.
Published: (2026)
by: Xu, Yang, et al.
Published: (2026)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
by: Obbad, Elyas, et al.
Published: (2024)
by: Obbad, Elyas, et al.
Published: (2024)
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization
by: Chan, Willy, et al.
Published: (2025)
by: Chan, Willy, et al.
Published: (2025)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models
by: Truong, Sang T., et al.
Published: (2024)
by: Truong, Sang T., et al.
Published: (2024)
On the Entropy Calibration of Language Models
by: Cao, Steven, et al.
Published: (2025)
by: Cao, Steven, et al.
Published: (2025)
The Use of a Large Language Model for Cyberbullying Detection
by: Ogunleye, Bayode, et al.
Published: (2024)
by: Ogunleye, Bayode, et al.
Published: (2024)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
by: Wang, Angelina, et al.
Published: (2025)
by: Wang, Angelina, et al.
Published: (2025)
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
by: Patel, Fagun, et al.
Published: (2025)
by: Patel, Fagun, et al.
Published: (2025)
Is Pre-training Truly Better Than Meta-Learning?
by: Miranda, Brando, et al.
Published: (2023)
by: Miranda, Brando, et al.
Published: (2023)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
by: Tang, Zeyu, et al.
Published: (2026)
by: Tang, Zeyu, et al.
Published: (2026)
Structured Prompts Improve Evaluation of Language Models
by: Aali, Asad, et al.
Published: (2025)
by: Aali, Asad, et al.
Published: (2025)
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
by: Dubois, Yann, et al.
Published: (2024)
by: Dubois, Yann, et al.
Published: (2024)
E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
by: Sadhuka, Shuvom, et al.
Published: (2025)
by: Sadhuka, Shuvom, et al.
Published: (2025)
Similar Items
-
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026) -
Extracting books from production language models
by: Ahmed, Ahmed, et al.
Published: (2026) -
Efficient Prediction of Pass@k Scaling in Large Language Models
by: Kazdan, Joshua, et al.
Published: (2025) -
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025) -
Logits are All We Need to Adapt Closed Models
by: Hiranandani, Gaurush, et al.
Published: (2025)