Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Nakkiran, Preetum, Bradley, Arwen, Goliński, Adam, Ndiaye, Eugene, Kirchhof, Michael, Williamson, Sinead |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Classifier-Free Guidance is a Predictor-Corrector
por: Bradley, Arwen, et al.
Publicado: (2024)
por: Bradley, Arwen, et al.
Publicado: (2024)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
por: Kirchhof, Michael, et al.
Publicado: (2025)
por: Kirchhof, Michael, et al.
Publicado: (2025)
Step-by-Step Diffusion: An Elementary Tutorial
por: Nakkiran, Preetum, et al.
Publicado: (2024)
por: Nakkiran, Preetum, et al.
Publicado: (2024)
BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
por: Choudhury, Deepro, et al.
Publicado: (2025)
por: Choudhury, Deepro, et al.
Publicado: (2025)
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
por: Santilli, Andrea, et al.
Publicado: (2025)
por: Santilli, Andrea, et al.
Publicado: (2025)
Vanishing Gradients in Reinforcement Finetuning of Language Models
por: Razin, Noam, et al.
Publicado: (2023)
por: Razin, Noam, et al.
Publicado: (2023)
Annotations Mitigate Post-Training Mode Collapse
por: Springer, Jacob Mitchell, et al.
Publicado: (2026)
por: Springer, Jacob Mitchell, et al.
Publicado: (2026)
Mechanisms of Projective Composition of Diffusion Models
por: Bradley, Arwen, et al.
Publicado: (2025)
por: Bradley, Arwen, et al.
Publicado: (2025)
Composition and Control with Distilled Energy Diffusion Models and Sequential Monte Carlo
por: Thornton, James, et al.
Publicado: (2025)
por: Thornton, James, et al.
Publicado: (2025)
Trace Length is a Simple Uncertainty Signal in Reasoning Models
por: Devic, Siddartha, et al.
Publicado: (2025)
por: Devic, Siddartha, et al.
Publicado: (2025)
The Geometries of Truth Are Orthogonal Across Tasks
por: Azizian, Waiss, et al.
Publicado: (2025)
por: Azizian, Waiss, et al.
Publicado: (2025)
Uncertainty Quantification for LLM Function-Calling
por: Ye, Zihuiwen, et al.
Publicado: (2026)
por: Ye, Zihuiwen, et al.
Publicado: (2026)
Calibration Across Layers: Understanding Calibration Evolution in LLMs
por: Joshi, Abhinav, et al.
Publicado: (2025)
por: Joshi, Abhinav, et al.
Publicado: (2025)
On the Calibration of Multilingual Question Answering LLMs
por: Yang, Yahan, et al.
Publicado: (2023)
por: Yang, Yahan, et al.
Publicado: (2023)
Task-Aware Calibration: Provably Optimal Decoding in LLMs
por: Tomov, Tim, et al.
Publicado: (2026)
por: Tomov, Tim, et al.
Publicado: (2026)
LLMs are not (consistently) Bayesian: Quantifying internal (in)consistencies of LLMs' probabilistic beliefs
por: Chen, Chacha, et al.
Publicado: (2026)
por: Chen, Chacha, et al.
Publicado: (2026)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
por: Jain, Neel, et al.
Publicado: (2024)
por: Jain, Neel, et al.
Publicado: (2024)
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment
por: Karim, Ahmed, et al.
Publicado: (2025)
por: Karim, Ahmed, et al.
Publicado: (2025)
Jailbreaking LLMs via Calibration
por: Lu, Yuxuan, et al.
Publicado: (2026)
por: Lu, Yuxuan, et al.
Publicado: (2026)
Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization
por: Zhang, Zheyuan, et al.
Publicado: (2026)
por: Zhang, Zheyuan, et al.
Publicado: (2026)
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
por: Gautam, Aayush, et al.
Publicado: (2025)
por: Gautam, Aayush, et al.
Publicado: (2025)
Calibrating Expressions of Certainty
por: Wang, Peiqi, et al.
Publicado: (2024)
por: Wang, Peiqi, et al.
Publicado: (2024)
When is Multicalibration Post-Processing Necessary?
por: Hansen, Dutch, et al.
Publicado: (2024)
por: Hansen, Dutch, et al.
Publicado: (2024)
Exploiting LLMs for Automatic Hypothesis Assessment via a Logit-Based Calibrated Prior
por: Gong, Yue, et al.
Publicado: (2025)
por: Gong, Yue, et al.
Publicado: (2025)
Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
por: Kirchhof, Michael, et al.
Publicado: (2025)
por: Kirchhof, Michael, et al.
Publicado: (2025)
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
por: Rofin, Mark, et al.
Publicado: (2026)
por: Rofin, Mark, et al.
Publicado: (2026)
Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering
por: Zhou, Han, et al.
Publicado: (2023)
por: Zhou, Han, et al.
Publicado: (2023)
Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause Frequencies
por: Liu, Terrance, et al.
Publicado: (2025)
por: Liu, Terrance, et al.
Publicado: (2025)
Local Mechanisms of Compositional Generalization in Conditional Diffusion
por: Bradley, Arwen
Publicado: (2025)
por: Bradley, Arwen
Publicado: (2025)
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
por: Paglieri, Davide, et al.
Publicado: (2024)
por: Paglieri, Davide, et al.
Publicado: (2024)
ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems
por: Zhou, Wenyong, et al.
Publicado: (2026)
por: Zhou, Wenyong, et al.
Publicado: (2026)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
por: Ouyang, Xu, et al.
Publicado: (2024)
por: Ouyang, Xu, et al.
Publicado: (2024)
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
por: Wang, Zhanliang, et al.
Publicado: (2026)
por: Wang, Zhanliang, et al.
Publicado: (2026)
Shielded Diffusion: Generating Novel and Diverse Images using Sparse Repellency
por: Kirchhof, Michael, et al.
Publicado: (2024)
por: Kirchhof, Michael, et al.
Publicado: (2024)
LLMs and the Madness of Crowds
por: Bradley, William F.
Publicado: (2024)
por: Bradley, William F.
Publicado: (2024)
LoCa: Logit Calibration for Knowledge Distillation
por: Yang, Runming, et al.
Publicado: (2024)
por: Yang, Runming, et al.
Publicado: (2024)
Geometry-Calibrated Conformal Abstention for Language Models
por: Xu, Rui, et al.
Publicado: (2026)
por: Xu, Rui, et al.
Publicado: (2026)
QA-Calibration of Language Model Confidence Scores
por: Manggala, Putra, et al.
Publicado: (2024)
por: Manggala, Putra, et al.
Publicado: (2024)
On the Importance of a Multi-Scale Calibration for Quantization
por: Son, Seungwoo, et al.
Publicado: (2026)
por: Son, Seungwoo, et al.
Publicado: (2026)
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
por: Dai, Hui, et al.
Publicado: (2026)
por: Dai, Hui, et al.
Publicado: (2026)
Ejemplares similares
-
Classifier-Free Guidance is a Predictor-Corrector
por: Bradley, Arwen, et al.
Publicado: (2024) -
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
por: Kirchhof, Michael, et al.
Publicado: (2025) -
Step-by-Step Diffusion: An Elementary Tutorial
por: Nakkiran, Preetum, et al.
Publicado: (2024) -
BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
por: Choudhury, Deepro, et al.
Publicado: (2025) -
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
por: Santilli, Andrea, et al.
Publicado: (2025)