Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges
Fuente:
arXiv
Salvato in:
| Autore principale: | Li, Yanran |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Finite-Calibration Regime Map for LLM Judge Panels
di: Zhu, Bin, et al.
Pubblicazione: (2026)
di: Zhu, Bin, et al.
Pubblicazione: (2026)
Efficient Inference for Noisy LLM-as-a-Judge Evaluation
di: Chen, Yiqun T, et al.
Pubblicazione: (2026)
di: Chen, Yiqun T, et al.
Pubblicazione: (2026)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment
di: Jung, Julie, et al.
Pubblicazione: (2025)
di: Jung, Julie, et al.
Pubblicazione: (2025)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
di: Zhou, Yukai, et al.
Pubblicazione: (2024)
di: Zhou, Yukai, et al.
Pubblicazione: (2024)
Segmenting Human-LLM Co-authored Text via Change Point Detection
di: Li, Mengchu, et al.
Pubblicazione: (2026)
di: Li, Mengchu, et al.
Pubblicazione: (2026)
Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
di: Hobelsberger, Christian, et al.
Pubblicazione: (2025)
di: Hobelsberger, Christian, et al.
Pubblicazione: (2025)
Optimal Estimation of Watermark Proportions in Hybrid AI-Human Texts
di: Li, Xiang, et al.
Pubblicazione: (2025)
di: Li, Xiang, et al.
Pubblicazione: (2025)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
di: Lacombe, Romain, et al.
Pubblicazione: (2025)
di: Lacombe, Romain, et al.
Pubblicazione: (2025)
End-To-End Causal Effect Estimation from Unstructured Natural Language Data
di: Dhawan, Nikita, et al.
Pubblicazione: (2024)
di: Dhawan, Nikita, et al.
Pubblicazione: (2024)
LIDS: LLM Summary Inference Under the Layered Lens
di: Park, Dylan, et al.
Pubblicazione: (2026)
di: Park, Dylan, et al.
Pubblicazione: (2026)
Don't Judge a Book by its Cover: Testing LLMs' Robustness Under Logical Obfuscation
di: Borah, Abhilekh, et al.
Pubblicazione: (2026)
di: Borah, Abhilekh, et al.
Pubblicazione: (2026)
Exploring Intra and Inter-language Consistency in Embeddings with ICA
di: Li, Rongzhi, et al.
Pubblicazione: (2024)
di: Li, Rongzhi, et al.
Pubblicazione: (2024)
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
di: Lin, Victoria, et al.
Pubblicazione: (2026)
di: Lin, Victoria, et al.
Pubblicazione: (2026)
Geological Inference from Textual Data using Word Embeddings
di: Linphrachaya, Nanmanas, et al.
Pubblicazione: (2025)
di: Linphrachaya, Nanmanas, et al.
Pubblicazione: (2025)
An Embedded Diachronic Sense Change Model with a Case Study from Ancient Greek
di: Zafar, Schyan, et al.
Pubblicazione: (2023)
di: Zafar, Schyan, et al.
Pubblicazione: (2023)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
di: Chew, Robert, et al.
Pubblicazione: (2026)
di: Chew, Robert, et al.
Pubblicazione: (2026)
MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment
di: Wang, Tianze, et al.
Pubblicazione: (2025)
di: Wang, Tianze, et al.
Pubblicazione: (2025)
Don't Touch My Diacritics
di: Gorman, Kyle, et al.
Pubblicazione: (2024)
di: Gorman, Kyle, et al.
Pubblicazione: (2024)
Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems
di: Landesberg, Eddie, et al.
Pubblicazione: (2025)
di: Landesberg, Eddie, et al.
Pubblicazione: (2025)
Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
di: Feng, Shangbin, et al.
Pubblicazione: (2024)
di: Feng, Shangbin, et al.
Pubblicazione: (2024)
Don't Pay Attention
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming
di: Kavumba, Pride, et al.
Pubblicazione: (2026)
di: Kavumba, Pride, et al.
Pubblicazione: (2026)
Bias and Uncertainty in LLM-as-a-Judge Estimation
di: Fiedler, James
Pubblicazione: (2026)
di: Fiedler, James
Pubblicazione: (2026)
A New Semisupervised Technique for Polarity Analysis using Masked Language Models
di: Watanabe, Kohei
Pubblicazione: (2026)
di: Watanabe, Kohei
Pubblicazione: (2026)
Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study
di: Okada, Kensuke, et al.
Pubblicazione: (2026)
di: Okada, Kensuke, et al.
Pubblicazione: (2026)
Syntax-Guided Diffusion Language Models with User-Integrated Personalization
di: Zhang, Ruqian, et al.
Pubblicazione: (2025)
di: Zhang, Ruqian, et al.
Pubblicazione: (2025)
Differential contributions of machine learning and statistical analysis to language and cognitive sciences
di: Sun, Kun, et al.
Pubblicazione: (2024)
di: Sun, Kun, et al.
Pubblicazione: (2024)
A chart review process aided by natural language processing and multi-wave adaptive sampling to expedite validation of code-based algorithms for large database studies
di: Wang, Shirley V, et al.
Pubblicazione: (2025)
di: Wang, Shirley V, et al.
Pubblicazione: (2025)
Aligning NLP Models with Target Population Perspectives using PAIR: Population-Aligned Instance Replication
di: Eckman, Stephanie, et al.
Pubblicazione: (2025)
di: Eckman, Stephanie, et al.
Pubblicazione: (2025)
Distributed Asymmetric Allocation: A Topic Model for Large Imbalanced Corpora in Social Sciences
di: Watanabe, Kohei
Pubblicazione: (2025)
di: Watanabe, Kohei
Pubblicazione: (2025)
Isolated Causal Effects of Natural Language
di: Lin, Victoria, et al.
Pubblicazione: (2024)
di: Lin, Victoria, et al.
Pubblicazione: (2024)
Transforming Sensitive Documents into Quantitative Data: An AI-Based Preprocessing Toolchain for Structured and Privacy-Conscious Analysis
di: Ledberg, Anders, et al.
Pubblicazione: (2025)
di: Ledberg, Anders, et al.
Pubblicazione: (2025)
Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models
di: Lyu, Y., et al.
Pubblicazione: (2025)
di: Lyu, Y., et al.
Pubblicazione: (2025)
A Position Paper on the Automatic Generation of Machine Learning Leaderboards
di: Timmer, Roelien C, et al.
Pubblicazione: (2025)
di: Timmer, Roelien C, et al.
Pubblicazione: (2025)
Maximizing Signal in Human-Model Preference Alignment
di: Kraus, Kelsey, et al.
Pubblicazione: (2025)
di: Kraus, Kelsey, et al.
Pubblicazione: (2025)
Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
di: Mao, Xin, et al.
Pubblicazione: (2024)
di: Mao, Xin, et al.
Pubblicazione: (2024)
Calibrating Pre-trained Language Classifiers on LLM-generated Noisy Labels via Iterative Refinement
di: Ye, Liqin, et al.
Pubblicazione: (2025)
di: Ye, Liqin, et al.
Pubblicazione: (2025)
Weighted Particle-Based Optimization for Efficient Generalized Posterior Calibration
di: Tanaka, Masahiro
Pubblicazione: (2024)
di: Tanaka, Masahiro
Pubblicazione: (2024)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
di: Zhou, Jin Peng, et al.
Pubblicazione: (2024)
di: Zhou, Jin Peng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Finite-Calibration Regime Map for LLM Judge Panels
di: Zhu, Bin, et al.
Pubblicazione: (2026) -
Efficient Inference for Noisy LLM-as-a-Judge Evaluation
di: Chen, Yiqun T, et al.
Pubblicazione: (2026) -
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
di: Moon, Jiwon, et al.
Pubblicazione: (2025) -
How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment
di: Jung, Julie, et al.
Pubblicazione: (2025) -
Don't Say No: Jailbreaking LLM by Suppressing Refusal
di: Zhou, Yukai, et al.
Pubblicazione: (2024)