"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
Fuente:
arXiv
Saved in:
| Main Author: | Hardy, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution
by: Hu, Zhengmian, et al.
Published: (2024)
by: Hu, Zhengmian, et al.
Published: (2024)
DeepScore: A Comprehensive Approach to Measuring Quality in AI-Generated Clinical Documentation
by: Oleson, Jon
Published: (2024)
by: Oleson, Jon
Published: (2024)
Limits of Large Language Models in Debating Humans
by: Flamino, James, et al.
Published: (2024)
by: Flamino, James, et al.
Published: (2024)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
by: Devunuri, Saipraneeth, et al.
Published: (2024)
by: Devunuri, Saipraneeth, et al.
Published: (2024)
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Domain-Shift-Aware Conformal Prediction for Large Language Models
by: Lin, Zhexiao, et al.
Published: (2025)
by: Lin, Zhexiao, et al.
Published: (2025)
RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning
by: Jin, Congyun, et al.
Published: (2024)
by: Jin, Congyun, et al.
Published: (2024)
From Traditional Taggers to LLMs: A Comparative Study of POS Tagging for Medieval Romance Languages
by: Schöffel, Matthias, et al.
Published: (2026)
by: Schöffel, Matthias, et al.
Published: (2026)
Improving LLM Leaderboards with Psychometrical Methodology
by: Federiakin, Denis
Published: (2025)
by: Federiakin, Denis
Published: (2025)
Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth
by: Davoudi, Seyed Pouyan Mousavi, et al.
Published: (2025)
by: Davoudi, Seyed Pouyan Mousavi, et al.
Published: (2025)
Metacognitive Myopia in Large Language Models
by: Scholten, Florian, et al.
Published: (2024)
by: Scholten, Florian, et al.
Published: (2024)
Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation
by: Foo, Leonardo Haw-Yang, et al.
Published: (2026)
by: Foo, Leonardo Haw-Yang, et al.
Published: (2026)
The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters Exam Performances
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
by: Tschisgale, Paul, et al.
Published: (2026)
by: Tschisgale, Paul, et al.
Published: (2026)
Uncertainty-Aware Adaptation of Large Language Models for Protein-Protein Interaction Analysis
by: Jantre, Sanket, et al.
Published: (2025)
by: Jantre, Sanket, et al.
Published: (2025)
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving
by: Kuzmanko, Jonathan
Published: (2025)
by: Kuzmanko, Jonathan
Published: (2025)
Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease
by: Dobbins, Nic, et al.
Published: (2025)
by: Dobbins, Nic, et al.
Published: (2025)
Augmented Risk Prediction for the Onset of Alzheimer's Disease from Electronic Health Records with Large Language Models
by: Wang, Jiankun, et al.
Published: (2024)
by: Wang, Jiankun, et al.
Published: (2024)
United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections
by: von der Heyde, Leah, et al.
Published: (2024)
by: von der Heyde, Leah, et al.
Published: (2024)
Language Models as Causal Effect Generators
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
Unified Representation of Genomic and Biomedical Concepts through Multi-Task, Multi-Source Contrastive Learning
by: Yuan, Hongyi, et al.
Published: (2024)
by: Yuan, Hongyi, et al.
Published: (2024)
A Rational Analysis of the Speech-to-Song Illusion
by: Marjieh, Raja, et al.
Published: (2024)
by: Marjieh, Raja, et al.
Published: (2024)
Beyond the Hype: Embeddings vs. Prompting for Multiclass Classification Tasks
by: Kokkodis, Marios, et al.
Published: (2025)
by: Kokkodis, Marios, et al.
Published: (2025)
ICE-ID: A Novel Historical Census Dataset for Longitudinal Identity Resolution
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
Language-Dependent Political Bias in AI: A Study of ChatGPT and Gemini
by: Yuksel, Dogus, et al.
Published: (2025)
by: Yuksel, Dogus, et al.
Published: (2025)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
by: Zhou, Cai, et al.
Published: (2026)
by: Zhou, Cai, et al.
Published: (2026)
Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
by: Hang, Haotian, et al.
Published: (2025)
by: Hang, Haotian, et al.
Published: (2025)
Crowdsourced Adaptive Surveys
by: Velez, Yamil
Published: (2024)
by: Velez, Yamil
Published: (2024)
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptoms
by: Vail, Alexandria K., et al.
Published: (2026)
by: Vail, Alexandria K., et al.
Published: (2026)
Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models
by: Nagori, Aditya, et al.
Published: (2025)
by: Nagori, Aditya, et al.
Published: (2025)
Removing Spurious Correlation from Neural Network Interpretations
by: Fotouhi, Milad, et al.
Published: (2024)
by: Fotouhi, Milad, et al.
Published: (2024)
Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
by: Miller, Evan
Published: (2024)
by: Miller, Evan
Published: (2024)
Mental Health Impacts of AI Companions: Triangulating Social Media Quasi-Experiments, User Perspectives, and Relational Theory
by: Yuan, Yunhao, et al.
Published: (2025)
by: Yuan, Yunhao, et al.
Published: (2025)
The Use of a Large Language Model for Cyberbullying Detection
by: Ogunleye, Bayode, et al.
Published: (2024)
by: Ogunleye, Bayode, et al.
Published: (2024)
Measuring Teaching with LLMs
by: Hardy, Michael
Published: (2025)
by: Hardy, Michael
Published: (2025)
Similar Items
-
A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution
by: Hu, Zhengmian, et al.
Published: (2024) -
DeepScore: A Comprehensive Approach to Measuring Quality in AI-Generated Clinical Documentation
by: Oleson, Jon
Published: (2024) -
Limits of Large Language Models in Debating Humans
by: Flamino, James, et al.
Published: (2024) -
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025) -
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
by: Wang, Hao, et al.
Published: (2026)