Fine-Tuning Language Models to Know What They Know

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Park, Sangjun, Meyerson, Elliot, Qiu, Xin, Miikkulainen, Risto
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918519724048384
author Park, Sangjun
Meyerson, Elliot
Qiu, Xin
Miikkulainen, Risto
author_facet Park, Sangjun
Meyerson, Elliot
Qiu, Xin
Miikkulainen, Risto
contents Evaluating true metacognition in Large Language Models (LLMs) is difficult due to biases and heuristics. This paper presents a framework to measure and enhance LLM metacognition while controlling for these biases. A measurement method using the $d'_{\rm type2}$ metric is established to isolate metacognitive ability. The Evolution Strategy for Metacognitive Alignment (ESMA) is proposed, demonstrating robust generalization across unseen datasets, languages, and newly acquired knowledge. Finally, parameter analysis reveals that these improvements are driven by a sparse set of parameters, offering new pathways for targeted metacognitive optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02605
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Fine-Tuning Language Models to Know What They Know
Park, Sangjun
Meyerson, Elliot
Qiu, Xin
Miikkulainen, Risto
Neural and Evolutionary Computing
Artificial Intelligence
Computation and Language
Neurons and Cognition
Evaluating true metacognition in Large Language Models (LLMs) is difficult due to biases and heuristics. This paper presents a framework to measure and enhance LLM metacognition while controlling for these biases. A measurement method using the $d'_{\rm type2}$ metric is established to isolate metacognitive ability. The Evolution Strategy for Metacognitive Alignment (ESMA) is proposed, demonstrating robust generalization across unseen datasets, languages, and newly acquired knowledge. Finally, parameter analysis reveals that these improvements are driven by a sparse set of parameters, offering new pathways for targeted metacognitive optimization.
title Fine-Tuning Language Models to Know What They Know
topic Neural and Evolutionary Computing
Artificial Intelligence
Computation and Language
Neurons and Cognition
url https://arxiv.org/abs/2602.02605