An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
Fuente:
arXiv
Saved in:
| Main Authors: | Pavlovic, Maja, Paun, Silviu, Poesio, Massimo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation
by: Pavlovic, Maja, et al.
Published: (2024)
by: Pavlovic, Maja, et al.
Published: (2024)
Understanding The Effect Of Temperature On Alignment With Human Opinions
by: Pavlovic, Maja, et al.
Published: (2024)
by: Pavlovic, Maja, et al.
Published: (2024)
Understanding Model Calibration -- A gentle introduction and visual exploration of calibration and the expected calibration error (ECE)
by: Pavlovic, Maja
Published: (2025)
by: Pavlovic, Maja
Published: (2025)
Uncertainty in Language Models: Assessment through Rank-Calibration
by: Huang, Xinmeng, et al.
Published: (2024)
by: Huang, Xinmeng, et al.
Published: (2024)
Extending Activation Steering to Broad Skills and Multiple Behaviours
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
by: Dai, Hui, et al.
Published: (2026)
by: Dai, Hui, et al.
Published: (2026)
Revisiting Uncertainty Estimation and Calibration of Large Language Models
by: Tao, Linwei, et al.
Published: (2025)
by: Tao, Linwei, et al.
Published: (2025)
Calibration vs Decision Making: Revisiting the Reliability Paradox in Unlearned Language Models
by: Shukla, Divyaksh, et al.
Published: (2026)
by: Shukla, Divyaksh, et al.
Published: (2026)
Geometric-Averaged Preference Optimization for Soft Preference Labels
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Improving Context-Aware Preference Modeling for Language Models
by: Pitis, Silviu, et al.
Published: (2024)
by: Pitis, Silviu, et al.
Published: (2024)
Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning
by: Krishnan, Ranganath, et al.
Published: (2024)
by: Krishnan, Ranganath, et al.
Published: (2024)
On Subjective Uncertainty Quantification and Calibration in Natural Language Generation
by: Wang, Ziyu, et al.
Published: (2024)
by: Wang, Ziyu, et al.
Published: (2024)
Perceptions of Linguistic Uncertainty by Language Models and Humans
by: Belem, Catarina G, et al.
Published: (2024)
by: Belem, Catarina G, et al.
Published: (2024)
ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
by: Baek, Jinheon, et al.
Published: (2024)
by: Baek, Jinheon, et al.
Published: (2024)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
by: Elesedy, Hayder, et al.
Published: (2024)
by: Elesedy, Hayder, et al.
Published: (2024)
Beyond Labels: Aligning Large Language Models with Human-like Reasoning
by: Kabir, Muhammad Rafsan, et al.
Published: (2024)
by: Kabir, Muhammad Rafsan, et al.
Published: (2024)
Multi-Agent Reasoning with Consistency Verification Improves Uncertainty Calibration in Medical MCQA
by: Martinez, John Ray B.
Published: (2026)
by: Martinez, John Ray B.
Published: (2026)
Synthetic vs. Gold: The Role of LLM Generated Labels and Data in Cyberbullying Detection
by: Kazemi, Arefeh, et al.
Published: (2025)
by: Kazemi, Arefeh, et al.
Published: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
by: Lee, Harrison, et al.
Published: (2023)
by: Lee, Harrison, et al.
Published: (2023)
Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering
by: Zhou, Han, et al.
Published: (2023)
by: Zhou, Han, et al.
Published: (2023)
Referential ambiguity and clarification requests: comparing human and LLM behaviour
by: Madge, Chris, et al.
Published: (2025)
by: Madge, Chris, et al.
Published: (2025)
Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
by: Li, Nan, et al.
Published: (2025)
by: Li, Nan, et al.
Published: (2025)
On the Entropy Calibration of Language Models
by: Cao, Steven, et al.
Published: (2025)
by: Cao, Steven, et al.
Published: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
by: Sahoo, Subramanyam
Published: (2026)
by: Sahoo, Subramanyam
Published: (2026)
A Study on the Calibration of In-context Learning
by: Zhang, Hanlin, et al.
Published: (2023)
by: Zhang, Hanlin, et al.
Published: (2023)
The Reliability Paradox: Exploring How Shortcut Learning Undermines Language Model Calibration
by: Bihani, Geetanjali, et al.
Published: (2024)
by: Bihani, Geetanjali, et al.
Published: (2024)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
by: Kossen, Jannik, et al.
Published: (2023)
by: Kossen, Jannik, et al.
Published: (2023)
MOSLD-Bench: Multilingual Open-Set Learning and Discovery Benchmark for Text Categorization
by: Costache, Adriana-Valentina, et al.
Published: (2026)
by: Costache, Adriana-Valentina, et al.
Published: (2026)
Report Cards: Qualitative Evaluation of Language Models Using Natural Language Summaries
by: Yang, Blair, et al.
Published: (2024)
by: Yang, Blair, et al.
Published: (2024)
Soft-Masked Diffusion Language Models
by: Hersche, Michael, et al.
Published: (2025)
by: Hersche, Michael, et al.
Published: (2025)
Calibrating Language Models with Adaptive Temperature Scaling
by: Xie, Johnathan, et al.
Published: (2024)
by: Xie, Johnathan, et al.
Published: (2024)
Soft Prompting for Unlearning in Large Language Models
by: Bhaila, Karuna, et al.
Published: (2024)
by: Bhaila, Karuna, et al.
Published: (2024)
Enhancing In-context Learning via Linear Probe Calibration
by: Abbas, Momin, et al.
Published: (2024)
by: Abbas, Momin, et al.
Published: (2024)
Improving Classification Performance With Human Feedback: Label a few, we label the rest
by: Vidra, Natan, et al.
Published: (2024)
by: Vidra, Natan, et al.
Published: (2024)
Instances and Labels: Hierarchy-aware Joint Supervised Contrastive Learning for Hierarchical Multi-Label Text Classification
by: Yu, Simon, et al.
Published: (2023)
by: Yu, Simon, et al.
Published: (2023)
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
Agents Thinking Fast and Slow: A Talker-Reasoner Architecture
by: Christakopoulou, Konstantina, et al.
Published: (2024)
by: Christakopoulou, Konstantina, et al.
Published: (2024)
Calibration Across Layers: Understanding Calibration Evolution in LLMs
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
On Calibration of Large Language Models: From Response To Capability
by: Yang, Sin-Han, et al.
Published: (2026)
by: Yang, Sin-Han, et al.
Published: (2026)
Similar Items
-
The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation
by: Pavlovic, Maja, et al.
Published: (2024) -
Understanding The Effect Of Temperature On Alignment With Human Opinions
by: Pavlovic, Maja, et al.
Published: (2024) -
Understanding Model Calibration -- A gentle introduction and visual exploration of calibration and the expected calibration error (ECE)
by: Pavlovic, Maja
Published: (2025) -
Uncertainty in Language Models: Assessment through Rank-Calibration
by: Huang, Xinmeng, et al.
Published: (2024) -
Extending Activation Steering to Broad Skills and Multiple Behaviours
by: van der Weij, Teun, et al.
Published: (2024)