GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
Fuente:
arXiv
Salvato in:
| Autori principali: | Sung, Yoo Yeon, Fleisig, Eve, Hou, Yu, Upadhyay, Ishan, Boyd-Graber, Jordan Lee |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
di: Gor, Maharshi, et al.
Pubblicazione: (2026)
di: Gor, Maharshi, et al.
Pubblicazione: (2026)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
di: Li, Zongxia, et al.
Pubblicazione: (2025)
di: Li, Zongxia, et al.
Pubblicazione: (2025)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
di: Srikanth, Neha, et al.
Pubblicazione: (2026)
di: Srikanth, Neha, et al.
Pubblicazione: (2026)
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
di: Fleisig, Eve, et al.
Pubblicazione: (2023)
di: Fleisig, Eve, et al.
Pubblicazione: (2023)
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
di: Hoyle, Alexander, et al.
Pubblicazione: (2025)
di: Hoyle, Alexander, et al.
Pubblicazione: (2025)
Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of Topic Models
di: Li, Zongxia, et al.
Pubblicazione: (2025)
di: Li, Zongxia, et al.
Pubblicazione: (2025)
First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models
di: Saphra, Naomi, et al.
Pubblicazione: (2023)
di: Saphra, Naomi, et al.
Pubblicazione: (2023)
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
di: Fleisig, Eve, et al.
Pubblicazione: (2024)
di: Fleisig, Eve, et al.
Pubblicazione: (2024)
Labeled Interactive Topic Models
di: Seelman, Kyle, et al.
Pubblicazione: (2023)
di: Seelman, Kyle, et al.
Pubblicazione: (2023)
SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
di: Mondal, Ishani, et al.
Pubblicazione: (2025)
di: Mondal, Ishani, et al.
Pubblicazione: (2025)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2025)
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2025)
Ghostbuster: Detecting Text Ghostwritten by Large Language Models
di: Verma, Vivek, et al.
Pubblicazione: (2023)
di: Verma, Vivek, et al.
Pubblicazione: (2023)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA
di: Kabir, Tasnim, et al.
Pubblicazione: (2026)
di: Kabir, Tasnim, et al.
Pubblicazione: (2026)
Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
di: Jaggi, Harbani, et al.
Pubblicazione: (2024)
di: Jaggi, Harbani, et al.
Pubblicazione: (2024)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
di: Li, Zongxia, et al.
Pubblicazione: (2024)
di: Li, Zongxia, et al.
Pubblicazione: (2024)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
di: Gu, Feng, et al.
Pubblicazione: (2025)
di: Gu, Feng, et al.
Pubblicazione: (2025)
NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
di: Shu, Matthew, et al.
Pubblicazione: (2024)
di: Shu, Matthew, et al.
Pubblicazione: (2024)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
di: Gor, Maharshi, et al.
Pubblicazione: (2024)
di: Gor, Maharshi, et al.
Pubblicazione: (2024)
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
di: Fleisig, Eve, et al.
Pubblicazione: (2025)
di: Fleisig, Eve, et al.
Pubblicazione: (2025)
SciDoc2Diagrammer-MAF: Towards Generation of Scientific Diagrams from Documents guided by Multi-Aspect Feedback Refinement
di: Mondal, Ishani, et al.
Pubblicazione: (2024)
di: Mondal, Ishani, et al.
Pubblicazione: (2024)
Standard Language Ideology in AI-Generated Language
di: Smith, Genevieve, et al.
Pubblicazione: (2024)
di: Smith, Genevieve, et al.
Pubblicazione: (2024)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
di: Fleisig, Eve, et al.
Pubblicazione: (2024)
di: Fleisig, Eve, et al.
Pubblicazione: (2024)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
di: Si, Chenglei, et al.
Pubblicazione: (2023)
di: Si, Chenglei, et al.
Pubblicazione: (2023)
TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation
di: Jeong, Yeil, et al.
Pubblicazione: (2026)
di: Jeong, Yeil, et al.
Pubblicazione: (2026)
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
di: Li, Zongxia, et al.
Pubblicazione: (2024)
di: Li, Zongxia, et al.
Pubblicazione: (2024)
CAF-Score: Calibrating CLAP with LALMs for Reference-free Audio Captioning Evaluation
di: Lee, Insung, et al.
Pubblicazione: (2026)
di: Lee, Insung, et al.
Pubblicazione: (2026)
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
di: Pu, Xiao, et al.
Pubblicazione: (2025)
di: Pu, Xiao, et al.
Pubblicazione: (2025)
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
di: Balepur, Nishant, et al.
Pubblicazione: (2025)
PapersPlease: A Benchmark for Evaluating Motivational Values of Large Language Models Based on ERG Theory
di: Myung, Junho, et al.
Pubblicazione: (2025)
di: Myung, Junho, et al.
Pubblicazione: (2025)
Calibrating Model-Based Evaluation Metrics for Summarization
di: Liu, Hongye, et al.
Pubblicazione: (2026)
di: Liu, Hongye, et al.
Pubblicazione: (2026)
SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models
di: Daynauth, Roland, et al.
Pubblicazione: (2025)
di: Daynauth, Roland, et al.
Pubblicazione: (2025)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
Task Calibration: Calibrating Large Language Models on Inference Tasks
di: Li, Yingjie, et al.
Pubblicazione: (2024)
di: Li, Yingjie, et al.
Pubblicazione: (2024)
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
di: Balepur, Nishant, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024) -
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
di: Gor, Maharshi, et al.
Pubblicazione: (2026) -
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
di: Sung, Yoo Yeon, et al.
Pubblicazione: (2024) -
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
di: Balepur, Nishant, et al.
Pubblicazione: (2025) -
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
di: Balepur, Nishant, et al.
Pubblicazione: (2025)